Adversarial Manipulation of Reinforcement Learning Policies in Autonomous Agents
Yonghong Huang, Shih-Han Wang · 2018
Machine learning classifiers are known to be vulnerable to intentional perturbation of inputs, namely adversarial examples. There are extensive studies of adversarial machine learning in the context of computer vision with high dimensional inputs. In this work, we show that adversarial attacks are also effective in targeting reinforcement learning policies based on low-dimensional sensory inputs in autonomous agent controls. Our results show that the two adversarial example crafting techniques significantly degrade test performance of the trained reinforcement learning policies. In addition, we compare adversarial examples with random noise on the effectiveness of the attacks. Furthermore, we study the importance of the input features, in terms of the impact of performance in adversarial attacks.