DatUS: Data-Driven Unsupervised Semantic Segmentation With Pretrained Self-Supervised Vision Transformer

Sonal Kumar, Arijit Sur, Rashmi Dutta Baruah · IEEE Transactions on Cognitive and Developmental Systems · 2024

Successive proposals of several self-supervised training schemes continue to emerge, taking one step closer to developing a universal foundation model. In this process, unsupervised downstream tasks are recognized as one of the evaluation methods to validate the quality of visual features learned with self-supervised training. However, unsupervised dense semantic segmentation has yet to be explored as a downstream task, which can utilize and evaluate the quality of semantic information introduced in patch-level feature representations during self-supervised training of vision transformers. Therefore, we propose a novel data-driven framework, DatUS, to perform unsupervised dense semantic segmentation as a downstream task. DatUS generates semantically consistent pseudo-segmentation masks for an unlabeled image dataset without using visual-prior or synchronized data. The experiment shows that the proposed framework achieves the highest MIoU (24.90) and Average F1 Score (36.3) by choosing DINOv2 and the highest Pixel Accuracy (62.18) by choosing DINO as the self-supervised training scheme on the training set of SUIM dataset. It also outperforms state-of-the-art methods for the unsupervised dense semantic segmentation task with 15.02% MIoU, 21.47% Pixel Accuracy, and 16.06% Average F1 Score on the validation set of SUIM dataset. It achieves a competitive level of accuracy for a large-scale COCO dataset.

Read the paper · More papers on PaperTik