Improve Robustness of Safe Reinforcement Learning Against Adversarial Attacks

Xiaoyan Wang, Yujuan Zhang, Lan Huang · 2024

Models trained with adversarial attack can be significantly improved stability and performance when faced with new uncertain environment. In this paper, we propose the robust training framework based on Wasserstein SA-MDP for safe RL, which combines ideas from reinforcement learning, adversarial examples generation, Wasserstein metric and robust training. This framework can improve the robustness of safe RL models against adversarial attacks by considering the dynamics of the adversarial interaction as a sequential decision-making process. We made the experiment in safety gym based on the proposed framework and report the experimental results.

Read the paper · More papers on PaperTik