Improvement of learning method of multi-agent system by sharing learning data
Qi Liu, Tomohiro Hayashida, Ichiro Nishizaki, Shinya Sekizaki · 2021
Deep reinforcement learning algorithms have been employed to solve the problems of multi-agent learning, especially the non-stationarity problem that reduces the learning efficiency. In this study, an architecture of improving the learning efficiency by using an adaptation of Actor-Critic method is proposed. It was clear that the convergence of learning became faster when the agents shared their experience data with the other agents at some rate. In two kinds of maze benchmarks, some numerical experiments was conducted to show the performance of the learning architecture. And the result of this numerical experiments showed that the agents achieved their purpose in the mazes by the proposed method. Even though the agents did not share all their data, they could be trained to take the optimal policies.