Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes

David M. Bossens · 2024

Robustness and safety constraints are key requirements for AI systems. The robust constrained Markov decision process is a recent task-modelling framework that incorporates behavioural constraints and robustness to reinforcement learning systems. Earlier work proposed the robust constrained policy gradient (RCPG) algorithm, which robustifies either the value or the constraint and updates the worst-case distribution through constrained optimisation on a sorted value list. Highlighting potential downsides of RCPG such as not robustifying the full constrained objective and the lack of incremental learning, this paper introduces algorithms to robustify the Lagrangian and to learn incrementally using gradient descent over an adversarial policy. A theoretical analysis derives the Lagrangian policy gradient for the policy optimisation and the Lagrangian adversarial policy gradient for the adversary optimisation. Empirical experiments injecting perturbations in inventory management and safe navigation tasks demonstrate the benefit of these modifications, and combining both modifications yields the best overall performance.

Read the paper · More papers on PaperTik