LSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image Segmentation
Qin Ma, Lingling Li, Xiaoqiang Lu, Licheng Jiao, Fang Liu, Wenping Ma, Xu Liu, Long Sun · IEEE Transactions on Geoscience and Remote Sensing · 2025
Referring Remote Sensing Image Segmentation (RRSIS) task aims to generate segmentation masks for target objects based on language descriptions. It requires precise localization while distinguishing between visually similar yet semantically distinct objects. Fusing vision-language features only during extraction causes information loss and semantic forgetting in the decoder, harming similar target distinction. Additionally, high-resolution remote sensing images present challenges, including complex backgrounds, diverse object scales, and intricate boundaries, limiting the effectiveness of previous methods. To address these issues, we propose the Long-term Semantic-guidance ConvFormer (LSCF) Network. First, we fuse multi-receptive-field local features extracted by the Multi-scale CoordConv (MCC) module with language-aware global features from the Cross-modal Attention (CA) module to obtain multi-modal representations. Second, the Sampling Attention (SA) module enables fine-grained vision context alignment under semantic guidance. Finally, the Global Language Fusion (GLF) module is incorporated in the decoder to maintain long-term vision-language alignment and mitigate semantic degradation. Experimental validation on the RefSegRS, RRSIS-D, and RISBench datasets demonstrates that LSCF achieves oIoU scores of 83.27%, 77.42%, and 74.88%, and mIoU scores of 77.44%, 64.25%, and 68.53%, respectively. On RefSegRS, LSCF surpasses the SOTA method FIANet by 5.53% (oIoU) and 9.58% (mIoU), while delivering competitive performance on RRSIS-D and RISBench. Code and experimental configurations will be released.