Cross-view geo-localization through consistency-guided transformer and adversarial feature alignment
Jie Chen, Yueshun He, Linglong Xiong, Yuankun Yang, Bo Fu · International Journal of Remote Sensing · 2026
Cross-view geo-localization aims to determine the geographic position of a query image by matching it with a geo-referenced remote sensing image database. In UAV-to-satellite localization, this task is particularly important for navigation, monitoring, and spatial perception under GNSS-denied or GNSS-degraded conditions. Existing methods have achieved considerable progress by learning discriminative image descriptors; however, they often provide limited explicit modelling of view-shared semantic regions between UAV and satellite images. As a result, the learned representations may be affected by viewpoint-induced appearance variations, background clutter, and view-specific responses.To address these limitations, this paper proposes a Consistency-Guided Transformer with Adversarial Feature Alignment (CGT-AFA) for UAV-to-satellite cross-view geo-localization. The proposed framework introduces learnable cross-view guidance vectors to aggregate shared structural cues through cross-attention, thereby guiding the encoder towards view-consistent semantic regions. In addition, a gradient reversal layer and a shared view discriminator are incorporated to suppress view-specific information and promote view-invariant representation learning. To further improve retrieval discrimination and training stability, a dynamic piecewise soft-margin triplet loss is designed by adaptively estimating the decision threshold from the batch-wise distributions of positive and negative distances. Experiments on University-1652 show that CGT-AFA achieves 92.87% Recall@1 and 93.85% AP in the Drone-to-Satellite retrieval task, and 94.67% Recall@1 and 92.33% AP in the Satellite-to-Drone retrieval task. Additional evaluation on SUES-200 further demonstrates the robustness of the proposed method under multi-height UAV acquisition conditions. In particular, CGT-AFA achieves AP values of 91.37%, 95.18%, 97.71%, and 98.65% in the Drone-to-Satellite task at 150 m, 200 m, 250 m, and 300 m, respectively. These results indicate that the proposed consistency-guided representation learning strategy can effectively improve UAV-to-satellite retrieval performance and provide competitive performance compared with recent cross-view geo-localization methods.