Exploiting Structure and Utilizing Agent-Centric Rewards to Promote Coordination in Large Multiagent Systems (Extended Abstract)
Chris HolmesParker, Adrian Agogino, Kagan Tumer · 2013
A goal within the field of multiagent systems is to achieve scaling to large systems involving hundreds or thousands of agents. In such systems the communication requirements for agents as well as the individual agents ’ ability to make decisions both play critical roles in performance. We take an incremental step towards improving scalability in such systems by introducing a novel algorithm that conglomer-ates three well-known existing techniques to address both agent communication requirements as well as decision mak-ing within large multiagent systems. In particular, we cou-ple a Factored-Action Factored Markov Decision Process (FA-FMDP) framework which exploits problem structure and establishes localized rewards for agents (reducing com-munication requirements) with reinforcement learning using agent-centric difference rewards which addresses agent deci-sion making and promotes coordination by addressing the structural credit assignment problem. We demonstrate our algorithms performance compared to two other popular re-ward techniques (global, local) with up to 10,000 agents.