Learning and Reasoning for Embodied Agents in Long-Horizon Decision-Making Tasks

liang ma · 2026

Embodied agents operating in real-world environments are required to perform long-horizon decision-making, which demands not only perception and action execution but also robust learning and reasoning over complex physical contexts. Recent advances in vision-language models (VLMs) have demonstrated promising capabilities in high-level reasoning and planning. However, their ability to support learning and reasoning in physically grounded, long-horizon tasks remains insufficiently understood, particularly in structured three-dimensional environments where spatial dependencies and physical constraints play a critical role. To systematically investigate this problem, this thesis introduces PhyBlock, a progressive benchmark designed to evaluate the learning and reasoning capabilities of VLM-based embodied agents in long-horizon decision-making tasks. Grounded in robotic 3D block assembly scenarios, PhyBlock formulates a structured evaluation framework that captures increasing levels of task complexity and decision dependencies. Specifically, the benchmark incorporates a four-level hierarchical assembly task design, together with complementary Visual Question Answering (VQA) samples, enabling comprehensive assessment of spatial reasoning, physical understanding, and the ability to handle sequential decision processes over extended horizons. The benchmark consists of 2,600 tasks, including 400 assembly tasks and 2,200 VQA tasks, and evaluates models along three key dimensions: partial completion, failure diagnosis, and planning robustness. Extensive experiments on 23 state-of-the-art VLMs provide a systematic analysis of their performance in long-horizon decision-making settings. The results show that, despite strong performance in low-level perception and short-horizon reasoning, current VLMs exhibit significant limitations in maintaining coherent reasoning and effective planning over extended decision horizons. Performance degrades substantially as task complexity increases, reflecting weaknesses in handling spatial dependencies and multi-step physical constraints. Further analysis suggests that these limitations stem from insufficiently grounded representations of physical structure and a lack of consistent reasoning across sequential decision steps. Overall, this work provides a unified evaluation framework for studying learning and reasoning in embodied agents, offering new insights into the challenges of long-horizon decision-making and highlighting directions for future research toward more physically grounded and robust intelligent systems.

Read the paper · More papers on PaperTik