CasTransNet: Transformer-Based Cascading U-Net Architecture for Medical Image Segmentation

Yong Wang, Lulu Zhang, Qingling Xia, Jianfei Pu, Yizhou Ding, Duoqian Miao · 2024

Most current medical image segmentation methods are based on transformers, whose self-attention mechanism, inspired by the human brain’s ability to focus attention, adeptly captures extensive relationships between pixels, thus efficiently modelling spatial context. Numerous studies have integrated transformer into U-shaped architectures for medical image segmentation, but have yet to effectively suppress irrelevant information from the encoder. To address this issue, we propose a novel Cascaded U-Net architecture enhanced with Transformers ( CasTransNet) to fully use low-level and high-level features for improved semantic information. The mixed attention module, serving as the core module of the encoder, captures local features and global contextual information through channel attention and spatial attention, respectively, achieving comprehensive feature extraction results. Spatial attention utilizes improved convolutional multi-head attention to effectively establish spatial contextual representations. The cascaded decoder effectively fuses features from skip connections and the previous decoder stage to bridge the semantic gap. Ultimately, multi-level outputs are obtained through the mixed attention module. Furthermore, introducing of multi-level loss functions enhances the flexibility of the overall training process. Our method achieves an average Dice score of 92.76% on the ACDC dataset. Additionally, our method achieves a Dice coefficient of 86.09% and an HD95 distance of 9.51 on the Synapse dataset. The experimental results demonstrate that our method outperforms most current medical image segmentation methods on both datasets. Our code is available at https://github.com/skyvlan-ai/CasTransNet/tree/master.

Read the paper · More papers on PaperTik