Discriminative Representation Learning for Remote Sensing Visual Question Answering
Yingda Lyu, Yan Han, Yu Liu, Haipeng Chen · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Recently, Remote Sensing Visual Question Answering (RSVQA) has attracted increasing attention from both academia and industry, which is the basis for understanding the underlying correspondence between remote sensing imagery and text descriptions. However, current methods are still insufficient in learning discriminative visual and textual representations for answer reasoning, mainly due to two reasons: (1) the remote sensing image environment is complex and changeable, and the target scales vary significantly, making it difficult to extract discriminative visual features; and (2) there is a lack of effective guidance from remote sensing domain knowledge to learn discriminative features. To this end, we propose a D iscriminative R epresentation L earning (DRL) method that includes two key strategies: visual feature enhancement and prior knowledge guidance. Specifically, we employ the Fourier transform to simulate the diverse visual environment and force the model to mine discriminative visual representations by imposing consistency constraints with the original features. In addition, we leverage the Remote Sensing Multimodal Large Language Model (RSMLLM) to generate captions rich in remote sensing domain-specific prior knowledge. These captions, derived from RSMLLM’s powerful knowledge integration and summarization capabilities, can then be compared and fused with visual representations to generate more discriminative representations, which are ultimately used for answer reasoning. Finally, recognizing that most existing RSVQA methods rely solely on static remote sensing images, we introduce RSVideoQA, a novel satellite video question answering dataset. This dataset is designed to facilitate the exploration of the rich spatio-temporal dynamics inherent in video sequences. Experimental results across three distinct datasets validate the effectiveness of our proposed method. Our dataset and code will be released at https://github.com/chill-han/DRL .