MultiSwin3D: A Swin Transformer-Based Multi-Task Model for 3D Medical Imaging
Fan Li, Lanyu Xu · 2025
In the field of medical imaging, AI-assisted techniques such as object detection, segmentation, and classification are widely employed to alleviate the workload of physicians and doctors. However, single-task models are predominantly used, overlooking the shared information across tasks. This oversight leads to inefficiencies in reallife applications. In this work, we propose MultiSwin3D, a novel Swin Transformer-based Multi-task model to address the limitations of single-task models by jointly performing 3D detection, segmentation, and classification. Our model uses a Swin Transformer as the shared encoder to generate multi-scale features, followed by CNN-based task-specific decoders. The proposed framework was evaluated on the BraTS 2018 and 2019 datasets, achieving promising results across all three tasks. Additionally, we compare the performance and efficiency of our multi-task model with that of single-task models. Our multi-task model significantly reduces computational costs and achieves faster inference speed while maintaining comparable performance to the single-task models, highlighting its efficiency advantage. To the best of our knowledge, this is the first work to leverage Swin Transformers for multi-task learning that simultaneously covers classification, segmentation, and detection tasks in 3D medical imaging, presenting its potential to enhance diagnostic processes. The code is available at https://github.com/fanlimua/MultiSwin.git.