Automatic Decomposition of Reward Machines for Decentralized Multiagent Reinforcement Learning

Sophia Smith, Cyrus Neary, Ufuk Topcu · 2023

In cooperative multiagent reinforcement learning (MARL), a team of agents learns to work together to complete a task. Centralized approaches to MARL quickly become intractable as the number of agents increases, necessitating decentralized learning algorithms which take advantage of task decompositions to train the agents individually. However, these task decompositions typically require careful human engineering. In this work, we develop algorithms to automatically decompose a team task into a collection of subtasks that can be used for decentralized reinforcement learning. We use reward machines-structured representations of reward functions-to encode team tasks and to automatically generate task decompositions that enforce the following three properties. 1) Task Consistency: We generate decompositions that are consistent with the team's task-if the agents individually learn to accomplish their subtasks, we guarantee that the composition of their learned behaviors will accomplish the original task. 2) Minimized Coordination: Inter-agent coordination during task execution can be costly. We minimize the coordination that's necessary to execute the decomposed tasks, which simplifies the decentralized learning problem by reducing each agent's interdependencies with its teammates. 3) Fairly Distributed: We maximize a weighted sum that balances the total utility of the agents and the fairness of the decomposition, which we define in terms of the distribution of assigned subtasks between the agents. Experimental results in three-agent and five-agent MARL tasks show the method's novel capabilities. The algorithm automatically generates task decompositions that are consistent with the team task, that reduce unnecessary coordination between the agents, and that take the agent's utility over subtasks into account. When used to define decentralized objectives, the generated task decompositions result in team policies that efficiently complete the task. Meanwhile, baseline decompositions yield policies that fail to complete the task.

Read the paper · More papers on PaperTik