Whole Semantic Sparse Coding Network for Remote Sensing Image–Text Retrieval

Chengyu Zheng, Qi Wen, Xiu Li, Chenxue Yang, Jie Nie, Yiyun Guo, Yuntao Qian, Zhiqiang Wei · IEEE Transactions on Geoscience and Remote Sensing · 2025

In recent years, cross-modal text-image retrieval in remote sensing has gained prominence as a research focus due to its potential to provide abundant, inclusive, and multi-perspective information. However, existing methods usually focus on salient features, but these salient features cannot describe the image or text completely, resulting in the loss of some important details and discriminable information. In this article, a Whole Semantic Sparse Coding Network (WSSCN) is proposed for remote sensing image-text retrieval to build a complete and reliable features description for further improving the performance of the retrieval model. Specifically, the WSSCN first designs a Whole Semantic Sparse Representation Coding (WSSRC) module by constructing a robust semantic library to transform the dense features matrix of the image and text into a whole semantic sparse matrix that enables multiple semantic decoupling and leads to a more precise and detailed features expression. Afterwards, the Intra- and Inter-Modal Consistency (IIMC) module is devised to improve the intra-modal and inter-modal consistency of the whole semantic sparse representation from different models. Finally, the Salient and Whole Semantic Adaptive Learning (SWSAL) loss is proposed to focus on salient information or whole semantic information by calculating an adjusted parameter. Quantitative and qualitative experiments are performed on four extensive datasets for cross-modal retrieval in remote sensing to showcase the notable effectiveness achieved by implementing a whole semantic sparse coding network.

Read the paper · More papers on PaperTik