Multi-agent Reinforcement Learning for Collaborative Behavior Optimization Based on Game Theory
Lihua Wang · Automation and Instrumentation · 2009
In order to overcome the problem of convergence in multi-agent system dynamic collaborative behavior op-timization,a novel algorithm named improved Pareto-Q learning based on general-sum game is presented,in which the global goal is the goal of multi-agent learning for local pareto joint behavior,and in which a distribution method for collective payoff based on common acceptability is also presented.The simulation result based on multi-robot be-havior coodindicates this algorithm is feasible and practicable.