A Projection-based Exploration Method for Multi-Agent Coordination
Hainan Tang, Juntao Liu, Zhenjie Wang, Ziwen Gao, You Li · 2024
In multi-agent reinforcement learning (MARL), states with high exploration value are difficult to be identified and coordinately visited, resulting in low learning efficiency. To this end, a projection-based exploration method for multi-agent coordination (PEMAC) is proposed. Goal states are selected using the count-based approach in the optimal projection space, of which the entropy of state distribution is maximal. Then, by reshaping the rewards in the replay buffer, agents are trained to visit those high-value states in a coordinated manner. In order to verify the effectiveness of the proposed method, comparative experiments are conducted in the multi-particle environment (MPE), in which dense-reward and sparse-reward settings are all both considered. Corresponding results suggest that PEMAC can effectively improve learning efficiency.