Output Feedback Reinforcement Learning Control for the Continuous-Time Linear Quadratic Regulator Problem

Syed Ali Asad Rizvi, Zongli Lin · 2018

In this paper, we present an output feedback reinforcement learning scheme to solve the LQR problem for continuous-time linear systems. The problem consists of finding the optimal feedback gain to achieve asymptotic stability without the knowledge of system dynamics and the information of the full state. An output feedback policy iteration algorithm is proposed that iteratively solves the ADP Bellman equation to find the optimal control parameters. Unlike the existing methods, the proposed scheme does not require any discrete approximation, and is not affected by the excitation noise bias. As a result, the need of a discounting factor, which has been a bottleneck in the past in achieving stability guarantee, is eliminated. The learned control parameters are optimal and match exactly the solution of the LQR Riccati equation. Simulation results show the effectiveness of the proposed scheme.

Read the paper · More papers on PaperTik