Understanding and analyzing approximate dynamic programming with gradient-based framework and direct heuristic dynamic programming.
Lei Yang · 2011
This research addresses control performance and convergence properties of some approximate dynamic programming (ADP) approaches. An overview of common ADP designs is first given, along with some illustrative examples. An existing framework of gradient-based system theoretic approach is then analyzed by examining the related gradient based policy iteration (GBPI) from an algorithmic perspective. The strong connection between the GBPI and policy iteration is revealed. A more intuitive implementation, based on a modified value iteration, of the GBPI is applied to an MDP problem, to provide some fundamental understanding of the learning and optimization features of the GBPI under the gradient-based framework. The direct heuristic dynamic programming (HDP) and three typical benchmark examples are used to introduce a unique analytical framework that can be applied to other learning control paradigms in addressing complex control problems. The sensitivity analysis and the linear quadratic regulator (LQR) design are employed. Applications of the direct HDP for nonlinear control problems beyond sensitivity analysis and the confines of LQR are developed and compared whenever appropriate to an LQR. The convergence proofs of the direct HDP with linear approximators under a linear LQR control setting are provided. A direct HDP design is further implemented in a tracking control setting to address a general nonlinear discrete-time system problem with filtered tracking error. The stability of the design is analyzed through Lyapunov approaches. Conditions are provided to assure the uniformly ultimate boundedness (UUB) of the closed-loop tracking error and the weight estimation error of both critic and action networks in the direct HDP design.