3D Abdominal Multi-Organ Segmentation Based on Mamba and Channel-Spatial Cross Attention

Yazhu Shi, Jing Zhou · 2024

Due to the complexity of the multi-organ structures presented in three-dimensional abdominal medical images, the varying tissue mediums, and the blurred edges, there is a need to construct a segmentation network capable of capturing long-range dependencies and global context awareness to enhance its segmentation accuracy. Traditional convolutional neural network architectures are limited by their receptive fields, the self-attention layers in Transformer are not scalable, and U-shaped segmentation networks often overlook the conceptual discrepancies between the encoder and decoder. Therefore, this paper proposes a 3D abdominal multi-organ segmentation network with Mamba and channel-spatial cross attention. The encoder and decoder stages utilize the DGFN-Mamba module to effectively capture long-distance dependencies and localized details, fully understanding the volumetric context of medical images. Enhancing the global context by employing 3D Channel-Spatial Cross Attention to grasp the dependencies of channel and spatial relationships within multi-scale encoder features. The skip connections have been improved to achieve a comprehensive fusion of global deep semantic and local shallow semantic information during the decoding process. Experimental results indicate that compared with other mainstream 3D medical image segmentation methods, our method shows an increase of 1.78% in the average DSC and 1.56% in the average NSD on the Abdomen CT dataset, compared to the U-MambaEnc method. On the Abdomen MRI dataset, our method shows an improvement of 3.37% in average DSC and 3.39% in average NSD over the nnU-Net method, thereby demonstrating the effectiveness and practicality of our approach.

Read the paper · More papers on PaperTik