Multi-Agent Reinforcement Learning via Directed Exploration Method
Yu Xie, Rongheng Lin, Hua Zou · 2022 2nd International Conference on Consumer Electronics and Computer Engineering (ICCECE) · 2022
In reinforcement learning, as a result of sparse reward feedback, it is difficult for agents to learn effective strategy in complex environment. Therefore, this paper proposes a multi-agent reinforcement learning algorithm called Exploration MATD3, which uses directed exploration method to improve the performance of multiple agents in environment with sparse rewards. The Exploration MATD3 algorithm is based on multi-agent twin delayed deep deterministic policy gradient (MA TD3), and uses reward shaping to drive agents to explore unknown states. Reward is combined of extrinsic reward and intrinsic reward. The extrinsic reward is original reward from environment, and intrinsic reward is calculated by the k-nearest neighbor states and random network distillation in each episode to reflect the novelty of new states. Evaluation is on Atari 2600 multiplayer games, and results verify that the proposed algorithm has better performance than MATD3 in environment with sparse rewards.