A Multi-Layer Embodied Data Infrastructure for Scalable World Modeling and Physical Intelligence: Toward Unified Human–Environment–Robot Perception for Embodied Foundation Models
Gu, Wilson · Zenodo (CERN European Organization for Nuclear Research) · 2025
Abstract The rapid evolution of embodied AI—combining perception, world modeling, actuation, and reasoning—has generated a pressing need for unified data systems capable of integrating multimodal experience across humans, robots, and complex environments. Current embodied datasets are fragmented, domain-specific, and insufficiently scalable, limiting the emergence of robust embodied foundation models (EFMs). This paper proposes a multi-layer embodied data infrastructure designed to support large-scale world modeling and physical intelligence, enabling EFMs to reason about, interact with, and learn from the physical world at human-level generality. We introduce a conceptual and operational blueprint spanning (1) the sensing–interaction layer, (2) the structured semantics layer, (3) the simulation–augmentation layer, and (4) the alignment–control layer. We further present a unified human–environment–robot perception framework, discuss its implications for embodied representation learning, and propose evaluation benchmarks toward holistic physical intelligence. Finally, we analyze system-level requirements for global-scale embodied data ecosystems, including privacy, governance, international collaboration, and multi-party interoperability. Our proposal aims to provide a foundation for next-generation embodied AI research and industrial deployment, bridging robotics, cognitive science, XR technologies, and foundation models. Key Contributions and Vision This vision paper addresses the critical data bottleneck in Embodied AI by proposing a novel, layered infrastructure for physical intelligence. Our core contributions include: A Four-Layer Data Architecture: A conceptual and operational blueprint for managing the complexity of embodied experience, inspired by network layering: Layer 1 (Sensing–Interaction): Collects and synchronizes raw, heterogeneous sensorimotor streams. Layer 2 (Structured Semantics): Standardizes the "meaning" of physical events through unified ontologies and relational graphs. Layer 3 (Simulation–Augmentation): Multiplies data via high-fidelity physics engines (MuJoCo, Isaac Gym) and photorealistic rendering (NeRFs, Gaussian Splatting), bridging the Sim-to-Real gap. Layer 4 (Alignment–Control): Ensures safety and human-compatibility through RLAIF and human preference models, providing policy feedback. Unified Human–Environment–Robot Triadic Perception: We propose unifying the perception loops of humans, the environment, and robots into a shared representation space. This tri-modal latent space is essential for cross-embodiment transfer learning and achieving general-purpose physical intelligence. Holistic Evaluation Benchmarks: We propose a new benchmark suite covering Embodied Perception, World Modeling, Physical Manipulation, Social–Physical Interaction, and Safety/Alignment to measure progress toward human-level physical intelligence. System-Level Blueprint: The paper concludes by analyzing the critical requirements for a global embodied data ecosystem, including data governance, international standardization, and ethical considerations (e.g., federated robot learning and multi-country compliance). This work provides the foundational architectural blueprint necessary to transition Embodied AI from fragmented, task-specific models to truly generalist Embodied Foundation Models (EFMs).