Transformer-based Visual Grounding with Inter-Modality Cross Attention

Yahong Rosa Zheng, Guo-Shiang Lin, K. W. Chang · 2025

Visual Grounding aims to locate a target region in an image according to a given natural language expression. This task plays a critical role in vision-language understanding and facilitates various downstream applications. In this paper, we propose a transformer-based visual grounding method. The proposed VG method comprises some parts: encoder module, inter-modality cross-attention module (IMCA), visual-language fusion module (VLF), and detection head. The IMCA mechanism is designed to enhance the semantic feature maps in the deeper layers. In addition, a multi-task learning strategy is adopted to raise the capability of feature extraction and alignment of the proposed method on the visual grounding task. To evaluate the performance of the proposed VG method, two common datasets, RefCOCO and RefCOCO+, are used. Experimental results demonstrate that the proposed method can not only locate the targets well but also outperform some existing methods for VG task.

Read the paper · More papers on PaperTik