Text-to-Image Activation for Open-Vocabulary Semantic Segmentation in Remote Sensing
Wenzhen Wang, Aoran Xiao, Wei Ping He, Hongyuan Zhu, Liang Xiao · IEEE Transactions on Geoscience and Remote Sensing · 2025
Open-vocabulary semantic segmentation in remote sensing aims to recognize arbitrary object categories from satellite imageries beyond a fixed label set, but its progress is constrained by the lack of fine-grained pixel-level supervision during training. Text-based weak supervision offers a scalable alternative by generating pseudo-labels from image-level annotations, yet in remote sensing scenarios it faces twofold challenges:1) incompleteness and unreliability of pseudo-labels, and 2) segmentation learning degradation induced by label noise. To address these issues, we propose a framework named Text-to-Image Activation for Open-Vocabulary Semantic Segmentation (T2ASeg), tailored for the spatial and semantic characteristics of satellite imagery. T2ASeg integrates two key modules to enhance pseudo-label quality and segmentation robustness. The Text-guided Adaptive Activation Module (TAAM) generates scale-aware and semantically consistent class activation maps, serving as reliable spatial and semantic cues for pseudo-label generation. The Multi-Scale Structure and Semantic Enhancement Module (MSSEM) combines structural modeling with class-aware semantic guidance, enabling the model to better capture fine-grained semantics and spatial structures under noisy supervision. By synergizing these modules, T2ASeg delivers more accurate and complete segmentation maps, even in the presence of weak and noisy labels. Extensive experiments on multiple remote sensing benchmarks show that T2ASeg outperforms state-of-the-art methods in both accuracy and boundary quality, demonstrating the value of combining textual guidance with adaptive spatial modeling for open-vocabulary segmentation. Code is available at github.com/WyneeWang/T2ASeg.