Safe Multi-Agent Reinforcement Learning via Dynamic Shielding
Yunbo Qiu, Yue Jin, Lebin Yu, Jian Wang, Xudong Zhang · 2024
Improving the safety of policies trained by multi-agent reinforcement learning (MARL) is an essential problem for practical utilization. Traditional methods for safe MARL either fail to improve safety during training process, or require strong prior knowledge about the specific task, such as human intervention, expert policy, and state transition model. However, in practical applications, the safety during training process is also important, and strong prior knowledge of the task is generally inaccessible. In this paper, we propose a novel algorithm Dynamic Shielding for MARL (DS-MARL), which utilizes simple prior knowledge including agents’ own motion model to provide a dynamic shield for MARL. DS-MARL aims to improve not only the safety of final policy, but also the safety of training process, without strong prior knowledge. Experimental results show that DS-MARL promotes the safety of both training process and final policy, and also increases success rate of final policy.