Refinement Of Reinforcement Learning Algorithms Guided By Counterexamples

Briti Gangopadhyay, Somi Vishnoi, Pallab Dasgupta · 2022

Reinforcement learning (RL) is a technique of learning optimal behaviors with high rewards through repeated interactions with an environment of interest. However, in many real-world applications, along with a high return, it is also required that the agent avoids undesirable states both during training and deployment. Reinforcement learning relies on maximizing a reward function to learn an optimal behaviour. Maximizing a reward does not guarantee safe behaviour. Also, it is important to account for parameter changes in the environment post-deployment which may lead a pre-trained policy to fail. In this paper, we explore a strategy that uses counterexamples to refine RL policies that have been tuned to maximize reward. We show that this method works for a wide variety of RL algorithms. We validate our methodology on Open AI Gym environments.

Read the paper · More papers on PaperTik