LOD: Latent Objective Discovery in Heterogeneous Multi-Agent Reinforcement Learning
Sun, Ke, Zhou Shi · arXiv (Cornell University) · 2025
In multi-agent reinforcement learning (MARL) settings, agents are often heterogeneous with diverse intrinsic utility functions. In this study, we introduce Latent Objective Discovery (LOD) as a framework for decomposing agents' distinct utilities into mixed underlying objectives in a low-dimensional latent space. Latent objective discovery aims to capture flexible and abstract coordination patterns among agents and each abstract objective is associated with a shared value function. During learning, all the learned values are aggregated by each agent according to their personalized weights learned in an adaptive manner. Therefore, agents inherit objective-level estimates for policy updates and value learning, enabling structured information sharing without requiring access to other agents' policies. Algorithmically, we incorporate our approach into Multi-Agent Proximal Policy Optimization (MAPPO) to exploit this structure. Theoretically, we establish convergence guarantees under linear function approximation within the actor-critic framework. Empirically, we extensively validate the advantage of introducing latent objective discovery in popular multi-agent testbeds with heterogeneous agents.