Online Optimal Control of Discrete-Time Systems Based on Globalized Dual Heuristic Programming with Eligibility Traces

Jun Ye, Yougang Bian, Biao Xu, Zhaobo Qin, Manjiang Hu · 2021

In this paper, an online adaptive dynamic programming (ADP) scheme that combines eligibility trace is presented for solving optimal control of discrete-time systems. In contrast with the forward view learning that requires to store additional vectors to update, the backward view learning of the proposed scheme employs online collected data and previous gradient information to update the neural network (NN) parameters at each step, which reduces the computational burden. In order to approximate the cost function more accurately to achieve a better policy improvement direction in the exploration process, the proposed algorithm introduces an independent costate network on the basis of the traditional HDP framework to approximate the costate function. By utilizing the costate as supplement information to estimate the cost function, the estimation accuracy has been greatly improved. Finally, two numerical examples are presented and the simulation results demonstrate the effectiveness and the advantage of computation efficiency of the presented method.

Read the paper · More papers on PaperTik