CLEAN rewards for improving multiagent coordination in the presence of exploration
Chris HolmesParker, Adrian Agogino, Kagan Tumer · 2013
In cooperative multiagent systems, coordinating the joint-actions of agents is difficult. One of the fundamental diffi-culties in such multiagent systems is the slow learning pro-cess where an agent may not only need to learn how to behave in a complex environment, but may also need to ac-count for the actions of the other learning agents. Here, the inability of agents to distinguish the true environmental dynamics from those caused by the stochastic exploratory actions of other agents creates noise on each agent’s reward signal. To address this, we introduce Coordinated Learn-ing without Exploratory Action Noise (CLEAN) rewards, which are agent-specific shaped rewards that effectively re-move such learning noise from each agent’s reward signal. We demonstrate their performance with up to 1000 agents in a standard congestion problem.