MixFormer: An End-To-End UAV Image-Matching Network Based on Mixed Attention Mechanism
Haitao Jia, Shiyi Xu, Jianhua Li, Zehang Lin, Wenbo Xu, Ren Li, Jian Li · 2024
Achieving autonomous navigation and positioning for UAV through computer vision technology has emerged as a key focus and hot topic in current research. The precise matching of UAV images is crucial for realizing this functionality. However, the intrinsic difference between satellite images and UAV images resulted in the sub-optimal performance of traditional matching algorithms. Although the depth-local feature matching of linear Transformers has demonstrated superior hierarchical matching capabilities in UAV image matching, the lack of fine local interactions between pixel tokens restricts their ability to extract highly accurate and localized correspondences. To address the limitations of existing linear Transformer matching algorithms, this paper proposes a novel method called MixFormer, a Mixed Attention Matching Transformer for UAV image feature matching. The algorithm comprises two main aspects: firstly, it introduces a hierarchical feature matching attention framework that operates attention at different scales, integrating global attention and local attention to achieve global context awareness and fine-grained matching, thereby enabling efficient and compact implementation. Secondly, it introduces a feature transformation module for transitional Convolutional Neural Network (CNN) to extract features with full receptive fields. Experimental results on benchmark tests confirm the effectiveness and efficiency of the proposed technique.