Visual and Textual Commonsense-Enhanced Layout Learning for Vision-and-Language Navigation
Fang Gao, Lei Shi, Jingfeng Tang, Jiabao Wang, Shaodong Li, Shengheng Ma, Jun Yu · IEEE Transactions on Automation Science and Engineering · 2025
In the Vision-and-Language Navigation (VLN) task, an agent must comprehend natural language instructions and execute precise navigation in complex environments. While significant progress has been made in the VLN field, the limited availability of navigation data hinders existing methods from fully learning the commonsense relationships between rooms and landmarks, which are crucial for environmental understanding and successful navigation. To address this issue, this work proposes a Visual and Textual Commonsense-Enhanced Layout Learning Model (ViTeC). We leverage the open-world knowledge embedded in large models by utilizing ChatGPT and BLIP-2 to provide commonsense information about environments. Specifically, BLIP-2 analyzes the room type corresponding to each panoramic image, while ChatGPT infers and provides knowledge about the most common landmarks within each room type. Moreover, to compensate for the agent’s lack of commonsense at the visual level, we employ Stable Diffusion to generate commonsense-based visual images, enhancing the agent’s visual perception. To ensure the agent effectively learns commonsense about the environment, we designed a Text Commonsense Layout Learning Module and a Visual Commonsense Layout Learning Module. These modules help the agent acquire environmental commonsense from both linguistic and visual perspectives, enabling it to utilize commonsense information effectively during navigation, thereby improving its environmental understanding and reasoning capabilities. Experimental results demonstrate that ViTeC achieves strong performance on REVERIE, R2R, and SOON datasets, exhibiting good generalization ability in complex environments. This validates the effectiveness of ViTeC in enhancing the agent’s environmental understanding and navigation capabilities.