ALVG: Training High-Quality Multi-modal Fusion Modules for Visual Grounding with Attention Loss

Sicheng Yang, Rongwei Yu · 2025

Visual grounding is the task of locating the relevant region in an image based on a textual description. Most existing methods rely on pre-trained visual and text encoder to extract features from images and text, which are fed into a fusion module to obtain fused features. To obtain high-quality fused features, researchers often design various complex multi-modal fusion modules. These fused features are processed using an encoder-decoder architecture to produce the final output. However, in the training process, the optimization objective is typically centered around the final predicted outputs, such as bounding boxes or instance segmentation masks, while the quality of the fused features often receives less direct attention. Therefore, the gradient usually needs to traverse a relatively long path when being backpropagated to the multi-modal fusion module, leading to diminished optimization effectiveness. In this paper, we propose ALVG, a simple and efficient visual grounding framework. We design a novel loss function to directly supervise the attention mechanism of multi-modal fusion modules, along with a simple but effective text-guided image enhancement module to complement it. The enhanced features are directly used for instance segmentation tasks, as well as object detection tasks. Experiments on six widely used Visual Grounding datasets, including RefCOCO/+/g, ReferIt, Flickr30K, and GRefCOCO, demonstrate the superiority of ALVG. Our method not only improves efficiency and convergence speed but also achieves state-of-the-art performance on these benchmarks.

Read the paper · More papers on PaperTik