Parameter-Efficient Reparameterization Tuning for Remote Sensing Image–Text Retrieval
Jian Bo Yang, Shengyang Li, Manqi Zhao · IEEE Transactions on Geoscience and Remote Sensing · 2025
Vision language models have been gradually adapted to various tasks of remote sensing domain with full fine-tuning paradigm,e.g., remote sensing image-text retrieval (RSITR). Superior performance enhancements have proven the powerful generalization and robustness of vision language models. However, full fine-tuning vision language models is resource-intensive and poses risks of overfitting. Moreover, existing RSITR methods usually assume that remote sensing images correspond to text captions one by one and utilize the bidirectional matching training objective, which is not aligned with evaluation benchmarks and real-world applications. To tackle the mentioned problems, we propose a novel Parameter-Efficient Reparameterization Tuning with Ranking and Matching (PERT-RaMa) framework, which effectively migrates the vision language model (i.e.CLIP) to RSITR task. To overcome the overfitting issue, we build a lightweight, plug-and-play module called Kronecker product for low-rank adaptation (KPLoRA). KPLoRA obtains higher intrinsic rank with fewer parameters. Furthermore, we design the Ranking and Matching (RaMa) training method that converts RSITR task into one-to-one matching and one-to-many ranking, which is aligned with current RSITR benchmarks and accelerates training speed through removal of unessential computations. Comprehensive experiments on three public RSITR benchmarks demonstrate that the effectiveness and efficiency of the proposed retrieval model. Our method outperforms full fine-tuning methods without CLIP by nearly 5-10%, and achieves comparable or superior retrieval capability than CLIP with full fine-tuning and parameter-efficient fine-tuning. Furthermore, our RaMa training method increases the training speed by 5x compared to current training method.