Ancrage multimodal pour l'apprentissage par renforcement pour la navigation et la manipulation robotiques utilisant une architecture insipré de l'espace de travail global

Léopold Maytié · INRIA a CCSD electronic archive server · 2025

The ability to perceive, integrate, and act upon information from multiple sensory modalities represents a fundamental characteristic of intelligent systems. While biological agents seamlessly combine visual, auditory, tactile, and other sensory inputs to navigate complex environments and make informed decisions, artificial systems have traditionally struggled to achieve comparable multimodal integration, particularly in reinforcement learning contexts where agents must learn optimal behaviors through environmental interaction. Multimodal perception presents unique computational challenges extending beyond simple sensor fusion. Each sensory modality carries distinct temporal dynamics, spatial resolutions, and semantic content that must be harmonized to create coherent world representations. Biological systems achieve this through sophisticated neural architectures enabling cross-modal plasticity, attention mechanisms, and hierarchical processing. The challenge is compounded by the need to learn meaningful correspondences between modalities without extensive supervision, requiring systems that can discover and exploit the underlying structure of multimodal data.Acting in multimodal environments introduces additional complexity, as agents must not only perceive through multiple channels but also coordinate actions simultaneously affecting different sensory modalities. Every action generates feedback across multiple sensory channels; for example, grasping an object provides visual, tactile, and auditory information simultaneously. This creates rich learning opportunities where agents can discover cross-modal relationships through behavioral exploration and learn policies seamlessly integrating responses across modalities, adapting their behavior based on the complex interplay between sensory channels and environmental dynamics. This thesis addresses the critical challenge of grounding multiple modalities for multimodal reinforcement learning by investigating how disparate sensory information can be effectively integrated through shared representations. Drawing inspiration from cognitive neuroscience, specifically Global Workspace Theory, we propose that multimodal integration can be achieved through a centralized workspace facilitating information broadcasting and selective attention across modalities. This approach provides a theoretical framework for understanding how human consciousness emerges from competition and cooperation between specialised subsystems, thus offering valuable insights for the design of multimodal artificial agents. Our approach is structured in several steps. We begin by implementing and validating a computational Global Workspace model for multimodal integration, then explore how the broadcast properties of this architecture can enhance reinforcement learning in multimodal environments. We further investigate combining our system with a world model enabling the agent to learn through imagination, and finally demonstrate practical applications in goal-conditioned robotic scenarios. Through this comprehensive exploration, we aim to establish Global Workspace Theory as both a biologically plausible and computationally effective framework for creating embodied agents capable of sophisticated multimodal understanding and action.

Read the paper · More papers on PaperTik