Directional-Semantic-Enhanced Visual Grounding for Remote Sensing Images

Guo Chang Hu, Bin Sun, Shutao Li, Chenglong Lei, Xiliang Li, Mingkui Tan · IEEE Transactions on Geoscience and Remote Sensing · 2025

Visual grounding for remote sensing images (RSVG) is a fundamental vision-language task, which aims to locate the objects referred to by the natural language expression from the RS images. Natural language expressions often rely heavily on detailed directional information to describe target objects. Therefore, the thorough and effective utilization of spatial directional information is crucial for accurately locating the referred objects within complex RS images. However, most existing RSVG methods fail to fully leverage directional information, leading to suboptimal outcomes. This paper introduces a novel directional semantic enhanced method for RSVG, dubbed DSEVG. Specifically, we propose a scale-adaptive language-guided interaction (SLI) module that derives scale-specific language features through a hierarchical language adaptation mechanism. These scale-specific language features guide the visual backbone to extract visual features highly relevant to referring expressions. Furthermore, we present a directional semantic enhancement (DSE) module that implicitly enhances directional semantics in referring expressions by leveraging visual spatial information. It also employs an explicit spatial semantic alignment loss to supervise this process, generating localization prompts with enhanced directional representations. These prompts are injected into the queries of each decoder layer to guide the model in effectively utilizing directional semantics. Experimental results on the DIOR-RSVG and OPT-RSVG benchmark datasets validate the effectiveness of the proposed method and demonstrate state-of-the-art performance. Code is available at: https://github.com/WH231203/DSEVG.

Read the paper · More papers on PaperTik