Multi-Scale Co-Attention Reinforced U-Net for Medical Image Segmentation
Yantao Song, Miao Zhang, Yunli Lu, Jiang Chang, Lu Chen · 2024
Medical image segmentation plays a crucial role in classifying anatomical structures for clinical analysis and intervention. However, the inherent variability and complexity in target structures make accurate segmentation tasks challenging. Traditional convolutional networks (e.g., U-Net) have limitations in feature extraction due to small receptive fields, while larger convolutional kernels (e.g., ConvNeXt) struggle with local context modeling and risk overfitting. In this paper, we propose a multi-scale reinforced U-Net that integrates the U-Net and ConvNeXt architectures as a dual-channel encoder, capturing both global and local features simultaneously through multi-scale convolutional kernels. Specifically, we designed a co-attention module that achieves effective feature fusion across different scales by embedding positional information into channel attention. Our approach demonstrates state-of-the-art performance on multiple medical image datasets, highlighting its effectiveness, superiority, and generalizability in addressing the complexities of medical image segmentation.