Multi-Agent Cooperative Exploration via Weighted Conditional Mutual Information in Sparse Reward Tasks
Jinyi Liu, Fangyu Li · 2025
Multi-agent Reinforcement Learning (MARL) algorithms have shown remarkable success in tasks requiring cooperation. However, in sparse reward scenarios, traditional MARL algorithms struggle to obtain useful guiding signals, which hinders effective learning. To enable MARL algorithms to learn optimal policies in sparse reward tasks, we propose a multi-agent cooperative exploration (MACE) method via weighted conditional mutual information. Specifically, we design an observation uncertainty estimation method using random network distillation (RND) to compute the uncertainty of the collected observations, thereby encouraging the agents to explore areas of the joint observation space with high uncertainty. Additionally, we develop an intrinsic reward that employs weighted conditional mutual information to measure the cumulative impact of an agent's actions on the exploration progress of others, overcoming the cooperation challenges posed by partial observability. Experimental results in multiple sparse reward environments show that MACE outperforms SOTA baseline MARL algorithms.