An Bi-Directional Sequence Inference Framework for Multi-Agent Reinforcement Learning

Wujun Xu, Kaiwen Xia, Shuai Wang, Xiaolei Zhou, Tian He, Lin Li · 2024

To improve the efficiency of multi-agent systems in the Internet of Things, multi-agent reinforcement learning (MARL) has been extensively studied. Although transformer-based models now treat decision-making in MARL as a sequence problem and achieve advanced performance, the action of agents often depends on previous states without direction in information sharing. To address this issue, we propose a Bi-directional Sequence Inference framework for Multi-Agent Reinforcement Learning (BSI-MARL), consisting of three components: Action-Observation Processing Module, Sequence Inference Module, and Policy Optimization Module. The Action-Observation Processing Module defines the state space, action space, and reward function for the agents. Based on the Transformer model, BSI-MARL designs an encoder-decoder module for bi-directional sequence inference to generate action sequences for multi-agent decisions. Additionally, the policy gradient optimization module broadens the action sampling window, improving training efficiency. The experiments conducted on Mujoco-Half Cheetah and Google Research Football demonstrate that BSI-MARL exhibits excellent performance in multi-agent decision-making and good stability and generalization ability.

Read the paper · More papers on PaperTik