Foundation Models for Robotic Tasks: Survey, Challenges and Future Directions
Muhammad Alam, M. A. Viraj J. Muthugala, Mohan Rajesh Elara · 2025
Intelligent robots capable of human-like interaction and understanding has been a long-standing goal of researchers. The advancements of large-scale Artificial Intelligence (AI) models such as Large Language Models (LLMs) have transformed the way robots interact with humans. More recently, Vision-Language Models (VLMs) and Vision-Language-Action models (VLAs) have further enhanced autonomous robotic capabilities. This review paper serves to provide a comprehensive survey of the integration of LLMs, VLMs and VLAs in robotics. We first introduce the fundamental concepts of language grounding in robotics, and its shift from classical approaches to modern largescale models. Then, we present the prominent model architectures and algorithms for robotic tasks. Lastly, we present the models in distinct sections, namely manipulation, navigation and localization, perception and scene understanding, decision making and planning, and human-robot interactions (HRI). We conclude the paper with the challenges and limitations of the applied models, and their potential future directions. This review paper would serve as a foundational reference for future embodied AI research.