YOLO-VG: Enhancing Multi-Stage Feature Interaction for Visual Grounding
Yang Liu, Le Jiang, Guoming Li, Xiaozhou Ye, Ye Ouyang · 2025
Visual Grounding refers to locating the target area described by a specified text in an image, which is a core technology that connects human language and the physical visual world. Existing methods often suffer from insufficient interaction between features from different modalities, resulting in weak semantic understanding. To address this limitation, we propose a novel Visual Grounding method, YOLO-VG, which enhances feature interaction across multiple stages within an open-vocabulary object detector, YOLO-World. First, we design a query-aware backbone network which allows the model to incorporate query information early during the feature extraction stage. Second, we introduce a multi-scale multi-modal feature interaction PAN module at the neck stage, which effectively improves the model's global perception and semantic understanding. Third, we propose an efficient data generation pipeline that leverages existing multimodal large language models and Visual Grounding models to generate high-quality data. Our YOLO-VG inherits the efficiency of YOLO-World and achieves conpetitive performance on several public Visual Grounding benchmarks.