Augmented Lagrangian Method for Instantaneously Constrained Reinforcement Learning Problems

Jingqi Li, David Fridovich-Keil, Somayeh Sojoudi, Claire Jennifer Tomlin · 2021 60th IEEE Conference on Decision and Control (CDC) · 2021

In this paper, we study the Instantaneously Constrained Reinforcement Learning (ICRL) problem, in which we are tasked to find a reward-maximizing policy while satisfying certain constraints at each time step. We first extend a result on the strong duality of Constrained Markov Decision Process (CMDP) in the literature and propose a sufficient condition for strong duality of the ICRL problem. Inspired by the Augmented Lagrangian Method in constrained optimization, we propose a new surrogate objective function for ICRL, which could be efficiently optimized by common policy-gradient based RL algorithms. We show theoretically that a feasible and optimal policy could be obtained by optimizing this surrogate function, under certain conditions related to the feasible policy set. Our empirical results on a tabular Markov Decision Process and two nonlinear optimal control problems, a constrained pendulum and a constrained half-cheetah, justify our analysis, and suggest that our method could promote safety during learning and converge in a smaller number of iterations compared to the existing algorithms.

Read the paper · More papers on PaperTik