Research on Reinforcement Learning Method Based on Multi-agent Dynamic GoalsResearch on sparse rewards based on dynamic target multi-agent reinforcement learning
Minghan Duan, Junsong Wang, Feng Jiang · 2023
Generally speaking, sparse reward algorithms are usually only applicable to static target tasks. In order to solve the sparse reward problem in multi-agent dynamic target scenarios, this paper proposes the Multi-agent Dynamic Hindsight Experience Replay (MADHER) algorithm. The main idea of MADHER algorithm is to fully utilize two failed trajectories to generate experience with rewards, in order to solve the problem of sample inefficiency caused by sparse rewards for dynamic targets. In order to meet the characteristics of post experience playback for MADHER dynamic targets, this paper designs an efficient experience playback buffer structure. This design can greatly save time and expenses. The research results indicate that the MADHER algorithm is an effective method for solving the dynamic multi-objective sparse reward problem of multi-agent systems. Compared to other methods, it has faster convergence speed and higher performance.