Ambiguity Resolution in Vision-and-Language Navigation with Large Language Models

Siyuan Wei, Chao Wang, Juntong Qi · 2024

Vision-and-Language Navigation (VLN) represents a novel approach that integrates visual perception and language understanding, thereby enabling autonomous agents to navigate complex environments based on natural language instructions. While traditional methods have enhanced navigation success rates, they continue to encounter obstacles in cross-modal alignment, generalization, and real-time decision-making. In order to address these issues, in this paper, we propose a VLN approach based on large language models (LLMs), utilising LLMs as planners. By vectorising both predefined and generated commands and employing k-nearest neighbours algorithm to identify the optimal match, the method resolves command mismatches and enhances system performance. The results of the experimental study demonstrate that our method significantly improves navigation success rates and efficiency in dynamic environments.

Read the paper · More papers on PaperTik