DSSINet: Dynamic token selection and spatial interaction framework for medical image segmentation
Xiaoyu Pan, Weilin Xu, Hanyun Xu · Alexandria Engineering Journal · 2025
Medical image segmentation plays a critical role in clinical diagnosis; however, most existing approaches struggle to balance local detail and global context, and they often fail to fully exploit multi-scale features. In this work, we introduce DSSINet , a novel framework that integrates dynamic token selection with spatial interaction to overcome these challenges. To alleviate the high computational cost and background interference inherent to Transformer-based pixel-level self-attention, we develop the Bi-Level Routing Attention (BLRA) module, which performs sparse routing at both region and token levels to focus exclusively on the top- k most relevant areas. To enhance the capture of fine-grained semantic details, we design the Feature Depth Fusion (FDF) module, establishing implicit spatial correlations between low- and high-level features to enrich multi-scale representations. We further incorporate the Adjacent Domain Perception Module (ADPM) to maintain intra-layer consistency and generate auxiliary edge maps that guide the decoding process. Finally, leveraging multi-graph inference, the Reverse Graph Decomposition (RGD) module iteratively reconstructs representations in a coarse-to-fine manner, yielding precisely refined boundaries. Extensive experiments on five public benchmarks demonstrate the effectiveness of DSSINet, achieving IoU scores of 86.21% (ISIC 2016), 85.25% (CVC-ClinicDB), 85.30% (Kvasir-SEG), 92.51% (Chest X-ray), and 94.26% (REFUGE2), outperforming state-of-the-art baselines. Moreover, DSSINet significantly reduces computational costs, with 82.2% fewer FLOPs and 13.1% fewer parameters compared to UNet plus.