Enhancing Large Language Models with RAG for Visual Language Navigation in Continuous Environments

Xiaoan Bao, Zhiqiang Lv, Biao Wu · Electronics · 2025

The task of Visual Language Navigation in Continuous Environments (VLN-CE) aims to enable agents to comprehend and execute human instructions in real-world environments. However, current methods frequently face challenges such as insufficient knowledge and difficulties in obstacle avoidance when performing VLN-CE tasks. To address these challenges, this paper proposes a navigation approach guided by a Retrieval Augmented Generation (RAG) large language model (RAGNav). By constructing a navigation knowledge base, we leverage RAG to enhance the input of LLM, enabling the generation of more precise navigation plans. Furthermore, we introduce a Prompt Enhanced Obstacle Avoidance strategy (PEOA) to improve the flexibility and robustness of agents in complex environments. Experimental results indicate that our method not only increases the navigation accuracy of agents but also enhances their obstacle avoidance capabilities, achieving a 2% and 2.32% increase in success rates on the public datasets R2R-CE and RxR-CE, respectively.

Read the paper · More papers on PaperTik