Multiagent Optimal Response Q-learning and its Convergence
Zhang Hua · 2004
Based on analysis of multiagent reinforcement learning, an agent optimal response learning rule is proposed provided the assumptions of opponents' policy. Q values have been proved to be convergent if opponents' policy satisfies certain restrictions, and experimental results of grid games are consistent with the convergence proof.