Transferring CLIP for visual grounding in remote sensing images

Linlin Liang, Yizhuo Quan, Chengbo Wang, Yuanfei Chang, Yanyou Qiao · International Journal of Digital Earth · 2025

Remote Sensing Visual Grounding (RSVG) task aims to localize specific objects in remote sensing (RS) images based on natural language queries and holds considerable potential for various applications. Existing approaches primarily rely on unimodal pre-trained encoders, leading to insufficient cross-modal information alignment. Moreover, these methods incur high computational costs for full fine-tuning of both visual and language encoders to achieve modality alignment, thereby constraining localization performance. Therefore, we propose a CLIP (Contrastive Language-Image Pretraining)-based remote sensing visual grounding framework, RSCLIPVG. RSCLIPVG employs a frozen CLIP model to extract visual and textual features, and we introduce a lightweight visual adapter to adapt visual representations, efficiently transferring the rich multimodal knowledge of CLIP to the RSVG scenario. Furthermore, a Multi-Level Collaborative Cross-modal Enhancement (MLCCME) module is developed to refine and integrate multi-level visual and textual features, enabling comprehensive cross-modal interaction and alignment. This effectively enhances the feature representation of target objects, thereby mitigating issues such as scale variations and cluttered backgrounds in remote sensing imagery. Experimental results on the DIOR-RSVG dataset indicate that our approach significantly outperforms previous methods. These research findings demonstrate the potential of CLIP in RSVG tasks, offering new solutions and perspectives for the RSVG field.

Read the paper · More papers on PaperTik