Leveraging Hybrid Referring Expressions for Referring Video Object Segmentation
Yan Li, Qiong Wang · 2024
Existing referring video object segmentation methods segment the target object referred by one referring expression or a simple group of multiple referring expressions by word concatenation in all video frames, which might still suffer from the issue of wrong object identification. To solve this problem, we propose a new transformer-based framework built upon hybrid referring expressions. To fully explore the multi-model fusion, three interaction methods are studied, and the early interaction shows the best alignment between discriminative video features and multi-language semantics. Moreover, two ways are explored to generate conditional queries in transformer-based methods, and the proposed word-group level method can better represent the target instance. Experiments on the large-scale Ref-YoutubeVOS benchmark show that the proposed model achieves 63.8% J & F, exceeding the state-of-the-art methods.