MAST: Multiagent Safe Transformer for Reinforcement Learning

Suhang Wei, Xianwei Wang, Xiang Feng, Huiqun Yu · IEEE Transactions on Cognitive and Developmental Systems · 2025

Safety remains a crucial challenge in the application of reinforcement learning. Multiagent safe reinforcement learning (MASRL) is an emerging field aiming to learn safe control policies that maximize cumulative rewards while satisfying the safety constraints of multiagent systems. However, existing research is limited and faces challenges such as environmental nonstationarity and the curse of dimensionality in action spaces, hindering the balance between performance and safety. To address these, this article proposes a multiagent safe reinforcement learning algorithm based on Transformer (MAST). The constrained optimization problem is transformed into an unconstrained one using the Lagrangian method. We also propose the multiagent total advantage decomposition theorem, establishing the connection between MASRL and sequence models. A Transformer-based framework is proposed, where a Transformer-based actor network generates joint actions in parallel during training while producing actions autoregressively during inference. Empirical evaluations on the safe multiagent MuJoCo (MAMuJoCo) benchmark show that MAST outperforms state-of-the-art algorithms by 13.06%. Our attention-based reward and safety critics achieve a 22.10% increase in rewards and an 83.58% reduction in safety costs. Additionally, the Transformer-based actor improves performance by 53.60%–111.93% compared to RNN-based methods.

Read the paper · More papers on PaperTik