A robust image geo-localization architecture with foundation vision model and feature mixing
Gaoshuang Huang, yang zhou, Xiaofei Hu, Chenglong Zhang, Luying Zhao, Wenjian Gan, Mingbo Hou · 2024
Obtaining the geographical location of images through image geo-localization technology is a highly significant task. However, existing image geo-localization methods struggle with accuracy under difficult conditions such as viewpoint changes, illumination variations, seasonal changes, and occlusions. To address these challenges, we proposed an image geo-localization architecture based on a foundation vision model and feature mixing. The architecture involves truncating and fine-tuning the foundation vision model DINOv2 to extract robust image features, which are then aggregated using an MLP-Mixer-based mix module to obtain robust and generalized image global features. This architecture significantly improves the accuracy of image geo-localization under difficult conditions. Experimental results demonstrate that the proposed architecture outperforms state-of-the-art (SOTA) methods in image geo-localization accuracy. Compared to SOTA methods, our architecture achieves accuracy improvements of 6.35%, 4.06%, and 6.30% on the test sets Tokyo 24/7, Nordland, and SF-XL-testv1, respectively, with viewpoint changes, illumination changes, season changes, and occlusions.