Perusmallit robottimanipulaatiossa

Veeti Roponen · Aaltodoc (Aalto University) · 2026

This thesis reviews the emergence and development of foundation models for robotic manipulation, with a particular focus on Vision-Language-Action (VLA) models. The review traces the progression from classical preprogrammed robotic control and early learning-based approaches to the current VLA paradigm, in which perception, language understanding, and action generation are unified within a single transformer-based architecture. Two dominant approaches to action generation are examined: discrete tokenization, which treats robot actions as language tokens, and continuous generation via diffusion or flow matching. Key models including RT-2, Octo, OpenVLA, π0, GR00T N1, and Gemini Robotics are analyzed and compared in terms of their architectures, training strategies, and capabilities. The thesis also examines the critical role of training data, highlighting the scarcity and fragmentation of real-world robot demonstrations and the emerging strategies, such as simulation-based trajectory generation and data augmentation, developed to address this bottleneck. Finally, several open problems are identified, including the absence of unified evaluation benchmarks, the challenges of continual learning and sim-to-real transfer, the integration of sensory modalities beyond vision, safety in human-centric environments, and the privacy implications of deploying cloud-dependent VLA models in private spaces. The review finds that while VLA models represent a significant step toward general-purpose robotic manipulation, substantial challenges in data, evaluation, safety, and privacy must be addressed before reliable real-world deployment becomes feasible.

Read the paper · More papers on PaperTik