VLMs bridging-enhanced Scene Semantic Reasoning Framework for Image-Text Matching

Yihua Gao, Junyu Chen, Mingyong Li · 2025

The main challenge in image-text matching lies in bridging the gap between visual and linguistic modalities for accurate cross-modal semantic alignment. While current mainstream methods enhance local feature interactions through region-word attention mechanisms, their isolated object modeling paradigm fails to capture deep semantic relationships, limiting fine-grained cross-modal reasoning capabilities. Research has explored structured relationship models like scene graphs; however, the visual modality faces inherent limitations: unlike text, which can build relationship graphs through lexical logic, visual scenes lack clear contextual semantics, resulting in issues such as ambiguous entity boundaries and distorted relationships that impede effective cross-modal alignment. This paper proposes a VLMs bridging-enhanced Scene Semantic Reasoning framework (VSSR). Based on modal characteristic differences, we construct a dual-path scene parsing framework: On the visual modeling, leveraging the strong semantic understanding capabilities of Vision-Language Models (VLMs) to generate dense scene semantic labels; On the textual modeling, fully exploiting the structured advantages of language by designing a graph attention network guided by relationship inductive bias to deeply mine implicit semantic associations between textual entities. To further bridge the modality gap, we create a multimodal collaborative representation space, using scene semantic labels as anchors to bridge the two modalities and achieve cross-modal knowledge transfer through joint semantic projection. While maintaining linear computational complexity, this architecture realizes fine-grained matching from scene-level (image-caption) to entity-level (image-entity) through relation-aware semantic modeling. Experiments on the Flickr30K and MS-COCO benchmark datasets demonstrate that VSSR outperforms existing state-of-the-art approaches in retrieval performance.

Read the paper · More papers on PaperTik