Luminate: Linguistic Understanding and Multi-Granularity Interaction for Video Object Segmentation
Rahul Tekchandani, R. Uma Maheshwari, Praful Hambarde, Satya Narayan Tazi, Santosh Kumar Vipparthi, Subrahmanyam Murala · 2024
Referring Video Object Segmentation (R-VOS) is a challenging task that involves segmenting objects in a video based on linguistic descriptions. In this paper, we introduce a novel multi-granularity referring video Object segmentation framework, termed as LUMINATE. The LUMINATE framework introduces a streamlined approach to cross-modal fusion. The proposed LUMINATE enhanced interaction between visual and textual modalities begins with cross-attention between the vision encoder’s query and the text encoder’s key-value pairs, and vice versa. The results are then concatenated with the respective queries of the vision and text encoders, fostering a comprehensive understanding of semantic relationships. The combined features are fed into the Transformer Encoder for further refinement and integration into the segmentation pipeline. Extensive experiments on benchmark datasets, including Ref-DAVIS, demonstrate that our proposed LUMINATE approach achieves better results than state-of-the-art methods in terms of Jaccard and F-measure evaluation metrics. Furthermore, the efficiency of our multi-object R-VOS variant is highlighted, achieving a threefold speed improvement while maintaining satisfactory segmentation performance. The proposed approach contributes to advancing the capabilities of R-VOS models, paving the way for improved multimodal reasoning and real-world applications.