CLIP-UG: CLIP-Driven Vision-Language Model for UAV-View Geo-Localization

Jiayi Wu, Guorui Feng · IEEE Transactions on Consumer Electronics · 2025

The objective of cross-view geo-localization is to match remote sensing images with corresponding images within the same scene, viewed from different perspectives. The advancement in Unmanned Aerial Vehicle (UAV) technology has underscored the significance of UAV-view geo-localization. Concurrently, the evolution of computer vision has led to the near saturation of performance of UAV-view geo-localization that solely relies on measuring the distance to image features. Therefore, additional modalities are required to further enrich context information. Recently, researchers discover that pretrained visual language models like CLIP demonstrate superior performance across various downstream tasks. Therefore, we propose a CLIP-driven method for UAV-view geo-localization. This method incorporates semantic branches into the classic contrast learning framework. To obtain the text description of the image from a different view, we employ two class-specific text description branches and merely fine-tune the class-specific category text. In addition, a text-image contrast loss is introduced to supervise the image encoder. A two-stage training strategy is also implemented to achieve better model convergence. Our experiments, conducted on two widely used datasets, University-1652 and SUES-200, illustrate the superiority and effectiveness of the proposed method.

Read the paper · More papers on PaperTik