MGFuseSeg: Attention-Guided Multi-Granularity Fusion for Medical Image Segmentation
Guoping Xu, Xuesong Leng, Chang Li, Xingwei He, Xinglong Wu · 2023
Convolutional Neural Networks (CNNs) have been widely used in medical image segmentation to efficiently develop computer-aided diagnosis systems. Due to the locality of convolutional operations, they can be used to extract fine-grained features, but with limitations in building global context and long-range spatial relationships. Recently, shifted window-based multi-layer perceptron (Swin-MLP) methods have demonstrated the ability to learn coarse-grained spatial features in a fixed-size window, but they are not well suited to dense-prediction tasks, such as medical image segmentation. To harmonize the strengths and mitigate the weaknesses of CNNs and SwinMLP in extracting features of varying granularity, we proposed two novel granularity fusion modules that use coarse-grained features to guide the fusion of fine-grained features based on the attention mechanism. Specifically, the first fusion module, named as BGFuse (Block Granularity Fuse), could fuse various scale block-grained features from Swin-MLP. The second fusion module, termed as LGFuse (Local Granularity Fuse), could fuse semantic coarse granularity information into fine granularity features. Equipped with these two fusion modules, we present a new attention-guided encoder-decoder network architecture (termed MGFuseSeg) for medical image segmentation. Without bells and whistles, the proposed MGFuseSeg significantly boosts the performance on three challenging segmentation benchmarks including Synapse, ACDC, and ISIC. Codes are available at https://github.com/apple1986/MGFuseSeg