Cooperative multi-agent reinforcement learning based on online heuristic extraction
Jun Wu, Xin Xu, Zhenping Sun, Yan Huang · 2011
Reinforcement learning has been an important technique for adaptive decision-making of multi-agent systems in uncertain environments. However, the curse of dimensionality in multi-agent reinforcement learning usually causes the slow learning convergence or even failure. In this paper, a novel Online Heuristics Extraction method, which can integrate the prior heuristic policy with a learned heuristic policy, is presented. The new method can be incorporated into a tabular or approximate cooperative multi-agent reinforcement learning algorithm so as to speed up the learning process. Simulation results on a cooperative learning task show that, with the new method, a much better learning convergence can be achieved.