SwinMas: Shifted Windows and Mask Unit Attention with auxiliary supervision for medical image segmentation
Mohammed Lawal, Dewei Yi · Biomedical Signal Processing and Control · 2025
Semantic segmentation models have been trying to improve out-of-domain generalization abilities by incorporating techniques like transfer learning, and few-shot learning among others. Segment Anything Models (SAMs) achieve promising zero-shot generalization by training on large datasets with millions of images and billions of masks. Existing approaches attempt to utilize SAMs by fine-tuning them on specialized images like medical images. This is done in an attempt to improve general segmentation and in-domain generalization performance. The limitations of this approach lie in the application of SAMs as a black box with the only interest being in the final output mask limiting information flow throughout the network, and SAMs ability to provide single-channel masks which do not allow for encoding class information limiting their applications in multi-class segmentation tasks. In this research, we present a segmentation model that utilizes SAMs to enrich low-level semantic information at the shifted-window style encoder part of the network as well as high-level semantic class information at the decoder part of the network through interface blocks. This network design of decoupling SAMs and strategically aggregating their features throughout the network leads to an outright increase in semantic segmentation performance and also allows for leveraging SAMs to improve multi-class segmentation performance while improving out-of-domain generalization ability. Experiments were conducted on multi-class wound segmentation datasets where the model outperforms other segmentation models. The model is also tested on binary segmentation tasks using the Kvasir-SEG polyp segmentation dataset and the Breast Ultrasound Segmentation dataset where the proposed method performs better than other segmentation models. The source code can be assessed on GitHub with this link: https://github.com/mantis180/SwinMas . • A hybrid sliding window and mask unit attention generalized image encoder that improves perspective during feature extraction. • A segmentation decoder that combines hierarchical semantic features with generalized representations to improve medical image segmentation performance. • An auxiliary segmentation head for guiding mask prediction. • Performance experiments with other segmentation models across different medical image tasks that demonstrate the advantages of the proposed design.