LLM-based reasoning for robotic planning: Robustness to task and environmental complexity

Filippo Favali, Lorenzo Sabattini, Valeria Villani · Robotics and Autonomous Systems · 2026

The transition of robotics toward social and general-purpose applications necessitates agents capable of natural language interaction with inexperienced users and real-time adaptation to unforeseen scenarios. Current research in the field integrates Large Language Models (LLMs) with traditional model-based controls to facilitate high-level reasoning toward task completion and intuitive cooperation. However, existing literature remains fragmented. Existing frameworks often prioritize agent-centric capabilities, overlooking the critical interplay between task complexity and environmental dynamics that dictates real-world reliability. This work presents a unified benchmarking framework for systematically evaluating LLM-based reasoning strategies for robotic planning under varying levels of task and environmental complexity. Through extensive benchmarking, our results demonstrate that Chain of Thought (CoT) combined with Self-Consistency is the most resilient methodology, ensuring high levels of safety and effectiveness across diverse scenarios. Notably, our analysis reveals that while model scale and environmental scale have marginal impacts on performance, the primary determinants of failure are task complexity and dynamic disturbances, such as perceptual occlusions and the unpredictable behavior of external agents.

Read the paper · More papers on PaperTik