Convergence and optimality of policy gradient primal-dual method for constrained Markov decision processes

Dongsheng Ding, Kaiqing Zhang, Tamer Başar, Mihailo R. Jovanović · 2022 American Control Conference (ACC) · 2022

We study constrained Markov decision processes with finite state and action spaces. The optimal solution of a discounted infinite-horizon optimal control problem is obtained using a Policy Gradient Primal-Dual (PG-PD) method without any policy parametrization. This method updates the primal variable via projected policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the lack of concavity of the constrained maximization problem in policy space, we exploit the underlying structure to provide non-asymptotic global convergence guarantees with sublinear rates in terms of both the optimality gap and the constraint violation. Furthermore, for a sample-based PG-PD algorithm, we quantify sample complexity and offer computational experiments to demonstrate the effectiveness of our results.

Read the paper · More papers on PaperTik