Semantic area optimization method with Dynamic Graph Neural Network for cross-modal image matching

Yanjia Tian, Yan Dong · Alexandria Engineering Journal · 2026

Cross-modal image matching is essential for autonomous driving and medical imaging, yet it confronts modal heterogeneity, redundant computation and dynamic interference. Existing methods suffer from a disconnection between semantic guidance and dynamic optimization. To solve this, we propose DG-SAMViT (Dynamic Graph-Enhanced Semantic Area Matching with Vision Transformer), a fully collaborative three-module model: ViT extracts global–local fusion features to unify cross-modal feature space; SAM filters high-confidence semantic regions via SOA (Semantic Object Area) and SIA (Semantic Intersection Area) to cut redundancy; DGNN optimizes dynamic matching relationships based on geometric consistency and feature similarity. Experiments on COCO 2017 and Flickr30K show DG-SAMViT outperforms nine benchmarks like LoFTR, achieving Flickr30K Recall@1 68.0%, COCO AUC@5°62.7% and mAP 42.3%, with robustness coefficient ≥ 0.74 in low-texture/occluded scenes. Ablation tests confirm the irreplaceable synergy of the three modules. Future work focuses on lightweight design and multimodal expansion. DG-SAMViT offers a high-precision solution for cross-modal retrieval and SLAM feature matching, boosting the practical application of cross-modal technologies.

Read the paper · More papers on PaperTik