Lightweight Visual-Semantic Token Transformer for 3D Medical Image Segmentation

Boxi Wu, Chenhui Yang · 2025

Transformer-based architectures have demonstrated outstanding performance in medical image segmentation, particularly in the analysis of 3D medical data, such as brain tumors. However, the computational complexity associated with these models, particularly due to the high number of parameters and large-scale voxel-level computations, has posed significant challenges for practical deployment. To overcome these limitations, this paper proposes a lightweight network framework designed for efficient segmentation. This framework integrates a visual-semantic token processing mechanism between a hierarchical encoder for multi-level semantic feature extraction and an all-Multilayer Perceptron (all-MLP) decoder. By mapping image features from the continuous voxel space to the discrete vector space, the proposed mechanism significantly reduces the need for computationally expensive voxel-level operations. Moreover, the Transformer's self-attention mechanism facilitates the capture of long-range dependencies across different tokens, enhancing the model's ability to detect correlations between distant voxels, which is crucial for improving segmentation accuracy. This approach is particularly effective for the segmentation of 3D brain tumors with multiple lesion areas. Experimental results demonstrate that the proposed framework, with just 5.4 million parameters, achieves competitive performance on the BraTS dataset, yielding Dice scores of 0.89 for Whole Tumor (WT), 0.81 for Tumor Core (TC), and 0.71 for Enhancing Tumor (ET), showcasing its efficacy and potential for clinical applications.

Read the paper · More papers on PaperTik