Where To Learn: Embodied Perception Learning Planned by Vision-Language Models
Juan Wang, Di Guo, Huaping Liu · IEEE Transactions on Cognitive and Developmental Systems · 2025
Embodied learning plays a crucial role in transferring self-learning agents to adapt to the environment. Existing embodied learning methods primarily rely on reinforcement learning (RL) exploration policy to collect inaccurate perceptual result samples for improving perceptual capabilities. However, RL-based exploration policies encounter several challenges such as the need for substantial data for training and the struggle to keep the diversity of the collected samples. In this article, we propose an embodied learning method that employs vision-language models (VLMs) as task planners, code planners, and path planners. Specifically, our method employs layout knowledge of the VLMs to decompose the embodied learning task into multiple subtasks and then convert each subtask into executable code, which will be executed to guide the agent to explore and collect the diverse samples in different types of rooms. Additionally, VLMs incorporate an external database to identify regions that enhance perceptual capabilities, and the agent will explore these poor perception regions to collect samples that can improve the perception performance. Experimental results demonstrate the effectiveness of our approach without the need for additional training.