Video object segmentation with self-supervised framework for an autonomous vehicle

Adwait Joshi, Hrishitaa Kurchania, Ashish Patwa, Koushik Sagar S, Ujwala Kshirsagar, Deepali Rahul Vora · IET conference proceedings. · 2023

Self-supervision framework targets to study useful representations from unlabeled pool of data as the unlabeled data [1] is available in abundance and getting good quality labeled data tends to be costlier and time taking. The vision transformers (ViT) [2] have achieved the present state of art performance for applying self-supervision in video segmentation tasks. The literature review discusses various video segmentation frameworks using contrastive methods [3]. Similarity algorithms have been applied to get optimized attention heads to discriminate foreground and background. It involves Knowledge Distillation (KD) [4] that helps student model learn from the teacher network, in which the instructor's logits are employed to train the student. It's most recognized for being a powerful model compression technique. Data augmentation [5] method aims to obtain the accuracy by providing various aspects of the same image. Our study underlines the importance of student-teacher encoder, multi-crop training and optimal use of small patches with ViTs replicated by proposed novel distillation model which mainly concerned about the avoidance of expected trivial solutions with the help of involved predictor model and centering process. It defines the synergy between autonomous vehicle and smart cities where perceiving precise environment can be achieved with the proposed model.

Read the paper · More papers on PaperTik