RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation With Foundation Models
Keyan Chen, Chenyang Liu, Bowen Chen, Jiafan Zhang, Zhengxia Zou, Zhenwei Shi · IEEE Transactions on Geoscience and Remote Sensing · 2025
Referring Remote Sensing Image Segmentation (RRSIS) facilitates flexible scene analysis by leveraging vision-language collaborative interpretation. However, conventional coupled frameworks typically perform pixel decoding after cross-modal fusion, conflating target localization (“where”) with boundary delineation (“how”). While recent decoupled approaches utilizing foundation models (e.g., SAM) attempt to separate these tasks, they predominantly rely on reductionist “box-to-mask” or “prompt-to-mask” pipelines. Such simple interfaces compress rich referring expressions into simplistic geometric constraints, severing semantic consistency and failing to leverage open-vocabulary visual-semantic alignment, particularly in multi-entity scenarios. To address these limitations, we proposeRSRefSeg 2, an framework that logically reformulates the workflow into sequential subtasks of “slack localization” and “refined segmentation.” Central to our approach is a novel cascaded second-order referring prompter, designed to construct a robust semantic bridge between CLIP’s open-world understanding and SAM’s segmentation capabilities. Specifically, we introduce an orthogonal subspace decomposition mechanism that separates text embeddings into complementary components. This enables implicit cascaded reasoning to isolate target attributes from background noise: first performing slack localization via cross-modal interaction to identify potential regions, and subsequently generating refined prompts to guide SAM’s precise delineation. Furthermore, we incorporate parameter-efficient tuning to align natural image priors with remote sensing domains. Extensive experiments on RefSegRS, RRSIS-D, and RISBench demonstrate that RSRefSeg 2 significantly outperforms state-of-the-art methods, achieving an approximate 3% improvement in gIoU while offering superior diagnostic interpretability. The code is available at: https://github.com/KyanChen/RSRefSeg2.