Optimal tracking control for discrete-time systems by model-free off-policy Q-learning approach
Jinna Li, Decheng Yuan, Zhengtao Ding · 2017
In this paper, a novel off-policy Q-learning is developed for solving linear quadratic tracking (LQT) problem of discrete-time (DT) systems, using only the measured data along the system trajectories. How to learn the optimal tracking control policy by off-policy approach and prove no bias of optimal solution probably caused by adding a probing noise to guarantee persistent excitation are two challenging issues when designing off-policy Q-learning algorithm focused in this paper. To this end, a behavior policy is introduced, and a novel off-policy Q-function based iterative Bellman equation is derived in terms of the relationship between Q function and value function. Consequently, an off-policy Q-learning algorithm is developed and its convergence as well as no bias are proved. Simulation results are given to verify the effectiveness of the proposed method.