A self-supervised, few-shot semantic segmentation study based on mobileViT model structure
Lei Zhou, Ruohan Gao, Jiaxiang Wang · 2023
Vision tasks such as semantic segmentation face challenges due to a lack of data and the need to reduce model parameters. In this study, we present a novel approach that unites several techniques, including the mobileViT model, and self-supervised training, to improve the performance of few-shot image segmentation. Our approach leverages unsupervised saliency estimation to obtain a pseudo-mask and contrast learning to enhance the feature representation. Furthermore, we introduce the MaskSplit method to divide up saliency masks into the artificial query and support sets. To train the few-shot segmentation model, we apply self-supervised meta-learning based on vision transformer. Importantly, our approach eliminates the need for manual segmentation of annotations during training, enabling a fully self-supervised learning setup. We evaluate the effectiveness of our approach on two popular benchmark datasets for few-shot semantic segmentation, namely Pascal-5i and COCO-20i. Our study demonstrates its effectiveness in small-sample semantic segmentation tasks.