MA-Net: A multi-attention segmentation network for anisotropic prostate magnetic resonance images
Zhenglin Yi, Jinbo Chen, Zhiyong Cai, Jiatong Xiao, Jinhui Liu, Haisu Liang, Chao Quan, Xiongbing Zu, Longxiang Wu, Jiao Hu · Engineering Applications of Artificial Intelligence · 2025
Prostate cancer (PCa) is the most common cancer among American men. Accurate and reliable segmentation of the prostate using magnetic resonance (MR) imaging is critical for the diagnosis and treatment of prostate cancer. Due to the inherent characteristics of prostate MR images, such as anisotropic resolution, large changes in shape and size, and low contrast between the gland and surrounding structures, the performance of traditional two-dimensional and three-dimensional based segmentation methods is limited. To address these issues, we propose a multi-attention-guided hierarchical U-shaped encoder-decoder network (MA-Net) to perform accurate and robust segmentation of prostate MR images. Specifically, an explicit decay attention (EDA) module is first used to introduce spatial distance prior knowledge into the encoding process and perform multi-scale token aggregation operations to generate rich feature representations. Secondly, a multiple cross-slice attention (MCSA) module is introduced in the skip connection. By applying multiple attention mechanisms to depth feature maps of different scales, all cross-slice information on anisotropic MR images is learned, thereby improving the segmentation performance at various positions of the prostate. Finally, the global-local interaction (GLI) module, which can interactively fuse long-range dependencies and local context information, is used to obtain a more accurate global-local feature representation and further improve the accuracy of prostate segmentation. Experimental results on three public prostate MR image datasets fully demonstrate the effectiveness and advancement of our method. In future work, we plan to integrate self-supervised pre-training, cross-modal adaptive fusion, and domain generalization methods to enhance the generalization ability of MA-Net for multi-modal and multi-center data.