Language-Guided Contextual Transformer Network for Referring Video Object Segmentation
Akash Yadav, Aditya Helode, Vishnu Prasad Verma, Mayank Lovanshi, Krishnarajanagar G. Srinivasa · 2024
Referring video object segmentation (RVOS) seeks to segment objects contained inside a video clip by creating a segmentation mask for the referred object by means of a natural language reference. Traditional methods relying on 3D convolutional neural networks (CNNs) often suffer from misalignment issues between spatial and temporal features in adjacent frames, leading to imprecise segmentation. Furthermore, the lack of global context often results in inaccurate segmentations, particularly in scenarios involving high temporal variations and occlusion. To tackle these issues, firstly, we introduce a multi-modal transformer based Cross-Modal Interaction (CMI) module to enrich both visual and linguistic contexts to enhance visual-linguistic alignment. Secondly, we propose a Global Context Object Clustering (GCOC) module aimed at recognizing and preserving the inter-frame correlations in challenging scenarios. Lastly, we integrate a Visual-Linguistic Semantic Loss (VLSL) mechanism to mitigate undesirable inductive biases introduced by unconstrained referring expressions during the feature aggregation process. We carry out comprehensive experiments on widely utilized RVOS datasets, A2D sentences, and Ref-Youtube-VOS. On Ref-Youtube-Vos,a$\mathcal{J}\&\mathcal{F}$score of 58.3 is achieved, which exceeds the previous methods by a significant margin. Moreover, the remarkable results of 50.1 mAP without image pretrain and 57.2 mAP with image pretrain on A2D sentences showcase the efficacy of our approach. Also, The mean and overall IoU scores surpass the results of existing state-of-the-art(SOTA) methods.