Offline Multi-Agent Reinforcement Learning in Custom Game Scenario
Shukla Indu, William Wilson, Althea C. Henslee, Haley R. Dozier · 2023
Offline reinforcement learning (RL) has garnered considerable attention in recent years due to its attractive capability of learning policies from offline datasets without environmental interactions. The goal of this research is to establish a proof of concept for offline Multi-Agent Reinforcement Learning (OMARL) in combat simulations. While MARL has made impressive progress, OMARL remains a relatively underexplored area. In this paper, we investigate the application of OMARL in a custom game environment and explore the transfer of this learning to a pre-collected dataset from combat simulations for a simplified scenario with no further environmental interactions. To gain a complete understanding of how OMARL works, we initially focus on a custom multi-agent grid environment for two agents. During the learning process, each agent needs to identify the environment dynamics and cooperate with the other agent. Agents are expected to learn a new policy from a single static dataset of previously collected data generated by random actions. Real-world mission planning is complex and involves a multi-agent system; therefore, our pre-collected dataset missions are restricted to motion and action planning in a discrete grid environment.