Co-LLaVA: Efficient Remote Sensing Visual Question Answering via Model Collaboration
Fan Liu, Wenwen Dai, Chuanyi Zhang, Jiale Zhu, Lin Xin Yao, Xin Li · Remote Sensing · 2025
Large vision language models (LVLMs) are built upon large language models (LLMs) and incorporate non-textual modalities; they can perform various multimodal tasks. Applying LVLMs in remote sensing (RS) visual question answering (VQA) tasks can take advantage of the powerful capabilities to promote the development of VQA in RS. However, due to the greater complexity of remote sensing images compared to natural images, general-domain LVLMs tend to perform poorly in RS scenarios and are prone to hallucination phenomena. Multi-agent debate for collaborative reasoning is commonly utilized to mitigate hallucination phenomena. Although this method is effective, it comes with a significant computational burden (e.g., high CPU/GPU demands and slow inference speed). To address these limitations, we propose Co-LLaVA, a model specifically designed for RS VQA tasks. Specifically, Co-LLaVA employs model collaboration between Large Language and Vision Assistant (LLaVA-v1.5) and Contrastive Captioners (CoCas). It combines LVLM with a lightweight generative model, reducing computational burden compared to multi-agent debate. Additionally, through high-dimensional multi-scale features and higher-resolution images, Co-LLaVA can enhance the perception of details in RS images. Experimental results demonstrate the significant performance improvements of our Co-LLaVA over existing LVLMs (e.g., Geochat, RSGPT) on multiple metrics of four RS VQA datasets (e.g., +3% over SkySenseGPT on “Rural/Urban” accuracy in the test set of RSVQA-LR dataset).