S2CCT: Self-Supervised Collaborative CNN-Transformer for Few-shot Medical Image Segmentation

Rongzhou Zhou, Ziqi Shu, Weixing Xie, Junfeng Yao, Qingqi Hong · 2024

Self-supervised pre-training followed by fine-tuning is a potent paradigm for few-shot learning, leveraging extensive unlabeled data with remarkable efficacy. Current self-supervised methods often lean towards Vision Transformers (ViTs) rather than CNN-Transformer hybrid architectures, which generally demonstrate superior performance. However, this reliance on ViTs can lead to poor perception of local features by the model. The challenge lies in designing a suitable proxy task for hybrid architectures like CNN-Transformers, which have significant structural differences. Additionally, the current organization of CNN-Transformer hybrid backbones is often sequential, hindering collaboration during pre-training and the acquisition of robust representations. To address these issues, we propose Self-Supervised Collaborative CNN-Transformer (S2CCT) for few-shot medical image segmentation. This framework introduces three innovative designs: (1) a composite proxy task based on image masking and image super-resolution tailored for CNN-Transformer hybrid architectures, enabling the backbone to acquire robust representations during pre-training that can be transferred to downstream tasks; (2) a parallel CNN-Transformer architecture that better attends to multi-scale features in images, making it more suitable for dense prediction tasks like image segmentation; (3) a sparse and dense feature fusion module to enhance collaboration between the two encoders. Experiments demonstrate that S2CCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks, i.e., ACDC and KiTs19. The code and pretrained models will be released soon.

Read the paper · More papers on PaperTik