Model-Free Reinforcement Learning for Fully Cooperative Multi-Agent Graphical Games
Qichao Zhang, Dongbin Zhao, Frank L. Lewis · 2018
In this paper, the optimal coordinated control problem for the homogeneous multi-agent graphical games with completely unknown dynamics is investigated. The off-policy reinforcement learning is proposed to approach the solution of the Hamilton-Jacobi equation under the framework of centralized training and decentralized execution. The actor-critic structure is adopted to learn the optimal control policies. Note that the critic network is centralized using the information from all the agents, and the parameter sharing scheme is adopted for the single actor network during the training process. For the execution process, the centralized critic network is not required, and only the trained actor network is used for each agent to obtain the control input based on its individual observation. For the implementation purpose, the neural network approximators with the actor-critic structure are constructed to approach the optimal centralized value function and the optimal policies for the multiagent graphical games. Finally, a simulation example is provided to demonstrate the effectiveness of the proposed algorithm.