3D medical image segmentation based on group mixing dual attention
Niu Guo, Yi Liu, Pengcheng Zhang, Jiaqi Kang, Zhiguo Gui, Lei Wang · Biomedical Signal Processing and Control · 2025
In recent years, Vision Transformers (ViTs) have become a significant research direction in 3D medical image segmentation due to their ability to capture long-range dependencies. However, traditional attention mechanisms in ViTs only model token–token correlations at a single granularity, struggling to effectively capture multi-granular interactions (e.g., token–group and group–group relationships). In this paper, we propose a novel 3D medical image segmentation method based on a Group-Mixing Dual Attention (GMDA) block, termed GMDA UNETR. The core innovation of GMDA UNETR lies in its GMDA module, which synergistically integrates multi-granularity feature interactions and a dual attention block to generate high-quality segmentation masks. Specifically, the GMDA module first employs a feature aggregation module to group queries, keys, and values into multiple granularities, capturing correlations at token–token, token–group, and group–group levels, thereby producing more discriminative feature representations. Subsequently, the dual attention module further refines the features. In this process, the channel attention submodule plays a crucial role. It computes inter channel correlation matrices. By doing so, it is able to enhance the representation of critical channels. Moreover, to maintain the local spatial details, it makes use of the channel compression and spatial excitation module. On the other hand, the spatial attention submodule takes a different approach. It employs a compressed projection strategy. This strategy is aimed at reducing the spatial dimensions of keys and values. As a result, it can efficiently capture long-range dependencies. It achieves this by computing attention matrices in the compressed space. The proposed model was systematically evaluated on the INSTANCE 2022 and ACDC datasets using five fold cross validation. For the INSTANCE 2022 task, the model achieved a mean Dice Similarity Coefficient (DSC) of 75.0% and 95% Hausdorff Distance (HD95) of 25.78 mm, demonstrating optimized boundary alignment and region coverage for complex lesions. On the ACDC dataset, the model attained DSC scores of 91.6% (right ventricle), 90.5% (myocardium), and 94.7% (left ventricle), with an overall mean DSC of 92.3%, validating its efficiency and robustness in cardiac anatomical structure segmentation. These results highlight the model’s potential for heterogeneous lesion segmentation and complex anatomical structure analysis, providing a viable technical framework for multi-task medical image segmentation.