SALA: Semantic alignment and localization alignment for visual grounding

Hongbing Li, Xinran Wang, Linyi Yang, Qi Li, Bo Xiao · Neurocomputing · 2025

Visual grounding focuses on establishing fine-grained alignment between specific regions and query expressions, which is increasingly essential as a cornerstone of visual intelligence. Despite recent success, existing methods often struggle with two main issues. Firstly, using independently pre-trained uni-modal encoders to extract expressive feature embeddings leads to a significant semantic gap between uni-modal features, hindering the effective interaction of visual-linguistic contexts. Secondly, the supervision provided by box annotations is inherently sparse and often underexploited, which limits the model’s ability to capture the fine-grained visual cues necessary to distinguish referent objects from the background, thereby leading to localization ambiguity. In this paper, we propose a Semantic Alignment and Localization Alignment (SALA) framework for visual grounding, which effectively bridges the cross- and uni-modal semantic gap and improves localization performance through patch-level and pixel-level alignment. This contributes to enhancing the consistency of representation before multimodal fusion, thereby improving the localization performance. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Codes will be made publicly available after acceptance.

Read the paper · More papers on PaperTik