Improving Cooperation among Self-Interested Reinforcement Learning Agents

Andrea Bonarini, Alessandro Lazaric, Marcello Restelli, J. Munoz De Cote · 2005

Abstract. In many Multi-Agent Systems (MAS), agents (even if selfinterested) need to cooperate in order to maximize their own utilities. Repeated play in social dilemmas (e.g., the iterated prisoner’s dilemma) is a challenging problem for learning algorithms, since the Nash Equilibrium (NE) solution, pursued by most of the existing algorithms, is often inappropriate in these settings, while Pareto efficient (PE) solutions guarantee a better outcome for each agent. In this paper we propose two principles (Change or Learn Fast and Change and Keep) aimed at improving cooperation among Q-learning agents in self-play. Using a nplayer and m-action version of the iterated prisoner’s dilemma, we show how a best-response learning algorithm, such as Q-learning, improved as proposed, can achieve better cooperative solutions in a shorter time. 1

Read the paper · More papers on PaperTik