Ambiguity Resolution in Vision-and-Language Navigation with Large Language Models
Siyuan Wei, Chao Wang, Juntong Qi · 2024
Vision-and-Language Navigation (VLN) represents a novel approach that integrates visual perception and language understanding, thereby enabling autonomous agents to navigate complex environments based on natural language instructions. While traditional methods have enhanced navigation success rates, they continue to encounter obstacles in cross-modal alignment, generalization, and real-time decision-making. In order to address these issues, in this paper, we propose a VLN approach based on large language models (LLMs), utilising LLMs as planners. By vectorising both predefined and generated commands and employing k-nearest neighbours algorithm to identify the optimal match, the method resolves command mismatches and enhances system performance. The results of the experimental study demonstrate that our method significantly improves navigation success rates and efficiency in dynamic environments.