Scaling Multi-Agent Reinforcement Learning via State Upsampling.

Luis Pimentel, Rohan Paleja, Zheyuan Wang, Esmaeil Seraj, James Ellis Grant Pagan, Matthew Gombolay · 2022

We consider the problem of scaling Multi-Agent Reinforcement Learning (MARL) algorithms toward larger environments and team sizes.While it is possible to learn a MARL-synthesized policy on these larger problems from scratch, training is difficult as the joint state-action space is much larger.Policy learning will require a large amount of experience (and associated training time) to reach a target performance.In this paper, we propose a transfer learning method that accelerates the training performance in such high-dimensional tasks with increased complexity.Our method upsamples an agent's state representation in a smaller, less challenging, source task in order to pre-train a target policy for a larger, more challenging, target task.By transferring the policy after pre-training and continuing MARL in the target domain, the information learned within the source task enables higher performance within the target task in significantly less time than training from scratch.As such, our method enables the scalability of coordination problems.Furthermore, as our method only changes the state representation of agents across tasks, it is agnostic to the policy's architecture and can be deployed across different MARL algorithms.We provide results showing that a policy trained under our method is able to achieve up to a 7.88× performance improvement under the same amount of training time, compared to a policy trained from scratch.Moreover, our method enables learning in difficult target task settings where training from scratch fails. SAND2022-8384CThis paper describes objective technical results and analysis.Any subjective views or opinions that might be expressed in the paper do not necessarily represent the views of the U.

Read the paper · More papers on PaperTik