Towards learning to best respond when losing control

Richard Klíma, Daniel Bloembergen, Michael Kaisers, Karl Tuyls · Centrum Wiskunde & Informatica (CWI), the national research institute for mathematics and computer science in the Netherlands · 2018

Many applications require control policies that are robust against rare but significant deviations, for example caused by hardware failures or adversarial attacks.In this article we propose a reinforcement learning method that uses a prior belief over potential compromisation to learn such a robust policy without requiring observations of the event during training.Our method can be easily combined with different standard reinforcement learning algorithms such as Q-learning and Expected SARSA.Experiments in a multi-agent coordination domain show that our method improves performance of standard algorithms, both with and without communication between the agents that occasionally lose control.

Read the paper · More papers on PaperTik