Reinforcement Learning for Stochastic Cooperative Multi-Agent Systems

Martin Lauer, Martin Riedmiller · 2004

We present a distributed variant of Q-learning that allows to learn the optimal cost-to-go function in stochastic cooperative multi-agent domains without communication between the agents. We motivate this approach from a theoretical standpoint by showing its relation to standard Q-learning. Also, a practical variant of the algorithm is proposed- called the ’reduced-lists algorithm ’-that deals with the problem of the combinatorial explosion of the number of joint actions. The principle behaviour of the algorithm is shown on a benchmark problem proposed by Boutilier and Claus. 1.

Read the paper · More papers on PaperTik