A reinforcement learning approach to problem solving for time-varying quadratic optimal control with unknown dynamics

Zhou-Yang Liu, Jiang-Hang Yu, Fei Wei, Qi Zhang, Xian-Li Wei · 2022 37th Youth Academic Annual Conference of Chinese Association of Automation (YAC) · 2022

The research of the linear quadratic regulator (LQR) problem of continuous-time linear systems with time-varying paramaters is carried out in this paper. As is known, the solution of the LQR problem is characterized by the induced differential Riccati equation (DRE), and the classical approaches to solving DRE problems usually presuppose a complete knowledge of system dynamics. However, this precondition does not alway s hold, as the system dynamics may be partially or completely unknown, or may vary over time. This paper aims to propose a model-free way to learn the optimal controller of continuous-time LTV systems from the date extracted from control inputs and system trajectories. More specifically, the process of solving the DRE is first converted to solving an iterative sequence of differential Lyapunov equations (DLE) via a policy iteration reinforcement learning mechanism. During each iteration, the Bézier control points technique is adopted to calculate the solution approximately associated with each DLE in a least-squares manner. The result given in this work shows that, with an appropriate choice of the Bézier function, the proposed off-policy reinforcement learning scheme guarantees convergence to an approximate optimal solution with an arbitrarily small neighborhood. Finally, a simulation example is used to illustrates the merits and effectiveness of the presented algorithm.

Read the paper · More papers on PaperTik