Optimal Response Learning and Its Convergence in Multiagent Domains

张化祥, 黄上腾, 乐嘉锦 · 东华大学学报:英文版 · 2005

In multiagent reinforcement learning, with different assumptions of the opponents' policies, an agent adopts quite different learning rules, and gets different learning performances. We prove that, in nultiagent domains, convergence of the Q values is guaranteed only when an agent behaves optimally and its opponents' strategies satisfy certain conditions, and an agent can get best learning performances when it adopts the same learning algorithm as that of its opponents.

Read the paper · More papers on PaperTik