Linguistic-Guided Feature Refinement For One-Stage Visual Grounding
Zhaoyi Zhang, Shaorong Xie · 2024
Visual grounding is a challenging multimodal task that aims to locate object instances based on provided natural language expressions. Existing methods extract visual feature maps independently using pre-trained backbones, and then engage in cross-modal interaction through a post-fusion approach. We argue that such a post-fusion mechanism fails to fully exploit the information contained in both modalities. Furthermore, the visual features extracted by the pre-trained backbone often diverge from those required for effective multimodal reasoning, as the backbone tends to be influenced by a priori knowledge derived from object recognition tasks. In this paper, we propose a semantically-modulated refinement framework that uses linguistic information from the outset to guide the extraction of visual features, thereby addressing the problem of inconsistency. Our framework comprises a language-modulated visual encoder, a refined feature pyramid network (FPN), and a grounding module. Extensive experiments on four common datasets demonstrate the effectiveness and advanced performance of our proposed method.