Learning against non-stationary opponents

Pablo Hernández-Leal, Enrique Muñoz de Cote, Luis Enrique Sucar · 2013

In multiagent systems, in order to make the best decisions, each agent has to take into account not only the strategy used by other agents but also how those strategies might change in the future. This is further exacerbated in open environments where strategies cannot be assumed to be rational. This paper studies repeated interactions between an agent and an opponent that changes its strategy over time (it is non-stationary). Our main contribution is a framework for fast learning changing non-stationary strategies. It uses decision trees to learn the most up to date opponent’s strategy. The agent’s learned model is continuously re-evaluated to assess strategy switches. Our method detects such strategy switches by measuring tree similarities. Aside from its fast learning process, decision trees can provide an easy interpretation of the opponent model which is useful for our approach. A second contribution is with regards to computing a policy against the opponent making the best use of the generated opponent model. For this, we propose a novel approach of transforming a decision tree into a Markov Decision Process. We evaluated the proposed approach in the iterated prisoner’s dilemma, outperforming state of the art algorithms in predictive accuracy when facing non-stationary strategies.

Read the paper · More papers on PaperTik