Learning Optimal Control Policies for Stochastic Systems with a Relaxed Bellman Operator.

Andrea Martinelli, John Lygeros · 2020

We introduce a relaxed version of the Bellman operator for q-functions and prove that it is still a monotone contraction mapping with a unique fixed point. In the spirit of the linear programming approach to approximate dynamic programming, we exploit the new operator to build a simplified linear program (LP) for q-functions. In the case of discrete-time stochastic linear systems with infinite state and action spaces, the solution of the LP preserves the minimizers of the optimal q-function. Therefore, even though the solution of the LP does not coincide with the optimal q-function, the policy we retrieve is the optimal one. The LP has fewer decision variables than existing programs, and we show how it can be employed together with reinforcement learning approaches when the dynamics is unknown.

Read the paper · More papers on PaperTik