Foundation model enhanced robot navigation

Sili Wang · DR-NTU (Nanyang Technological University) · 2026

Visual Place Recognition (VPR) is a crucial component for autonomous mobile robot navigation, traditionally relying on deep metric learning to retrieve geo-tagged database images. However, standard metric learning objectives depend on a strict binary definition of positive samples. By treating all images within a physical distance threshold equally, these methods neglect the continuous nature of spatial feature distributions. This extreme compactness frequently causes severe feature collapse, stripping the model of fine-grained instance-level discriminability. To address this limitation, this dissertation proposes a distance-aware training paradigm centered on the Hierarchical Multi-Similarity (Hier-MS) Loss. The Hier-MS Loss constructs a pseudo-spatial hierarchy utilizing data augmentation. It explicitly preserves intra-class variance by defining highly perturbed augmented views as strict positive matches, while treating other distinct original images from the same location as soft negatives. To deploy modern Vision Foundation Models (VFMs) under constrained computational resources, the proposed architecture integrates a DINOv3 backbone adapted via Low-Rank Adaptation (LoRA), followed by a MixVPR global feature aggregator. Extensive experiments are conducted across five diverse datasets, encompassing large-scale outdoor urban environments, complex indoor scenarios, and a proprietary industrial setting. The evaluations demonstrate the effectiveness of the proposed framework. Compared to standard metric learning baselines, the Hier-MS Loss consistently improves Top-1 retrieval accuracy (Recall@1) while yielding lower mean and median absolute localization errors. These results confirm that the proposed method not only maintains robust binary retrieval performance but also significantly enhances fine-grained physical localization precision, providing a highly discriminative foundation for visual spatial reasoning.

Read the paper · More papers on PaperTik