Attentive Multi-Kernel Feature Aggregation Network for Cross-View Geo-Localization
Shuheng Huang, Deyong Wu, Jinliang Lin, Lei Peng, Zhiming Luo · 2025
Cross-view geo-localization, which aims to match images of the same scene captured from diverse viewpoints from drone and satellite, presents a persistent challenge due to significant geometric distortions and appearance variations. Existing methods lack a comprehensive exploration and dynamic integration of spatial and channel attention mechanisms, while primarily focusing on extracting single-scale features. The former results in the model failing to focus on key regions of the feature map, while the latter may result in the inability to capture information at different scales. In this paper, we propose a novel Attentive Multi-kernel Feature Aggregation (AMFA) Network that incorporates a Synergistic Attention (SA) Module and a Multi-kernel Inception (MI) Module, which effectively addresses the challenges posed by significant variations in target building regions and contextual diversity in cross-view tasks. The SA Module adaptively fuses channel and spatial information to focus on the most discriminative regions within the feature maps. Building upon this, the MI Module extracts multi-scale features through parallel convolutional kernels, enabling a more comprehensive scene representation. Experiments on the University-1652 and SUES-200 datasets show that our method achieves state-of-the-art performance in cross-view geo-localization tasks.