Exploring Multi-Modal Representation Learning for Embodied Agents
Zhenfei Yin · The Sydney eScholarship Repository (The University of Sydney) · 2026
Building intelligent embodied agents requires representations that support grounding, prediction, reasoning, and action. Existing systems show a gap between semantic understanding and embodied competence: they describe scenes and generate futures, yet fail to localize objects, reason over geometry, coordinate agents, or execute actions reliably. Representation learning is a fundamental bottleneck for embodied intelligence. This thesis investigates multi-modal representation learning for embodied agents. We study instruction-tuned multi-modal models across images, point clouds, and videos, showing semantic alignment emerges earlier than geometric grounding. We introduce evaluation protocols probing robustness, hallucination, and failure modes, and explore parameter-efficient adaptation strategies preserving grounding across tasks. We examine predictive models as world representations for embodied generalization. Visual realism alone is insufficient; effective world models must integrate multi-modal structure and physical constraints. Through evaluation combining perceptual assessment and embodied execution, models grounded in multi-modal representations—linking vision, geometry, language, and dynamics—generalize more reliably than unguided approaches. Long-horizon reasoning emerges from systems organizing predictive and control skills into adaptive multi-agent structures. Reward-driven self-organizing frameworks enable task decomposition and coordination. Structured constraints and hierarchical supervision show compositional multi-modal representations enable safer, more generalizable embodied behavior. Embodied intelligence is fundamentally a representation and system design problem. Competence emerges when multi-modal representations are task-conditioned, physically grounded, and compositional. These findings provide guidance for building general-purpose embodied agents.