The Adaptive Dynamic Programming Theorem

J. Murray, Chadwick J. Cox, Richard E. Saeks · Birkhäuser Boston eBooks · 2003

The centerpiece of the theory of dynamic programming is the HamiltonJacobi-Bellman (HJB) equation, which can be used to solve for the optimal cost functional V o for a nonlinear optimal control problem, while one can solve a second partial differential equation for the corresponding optimal control law k o .Although the direct solution of the HJB equation is computationally untenable, the HJB equation and the relationship between V o and k o serves as the basis for the adaptive dynamic programming algorithm. Here, one starts with an initial cost functional and stabilizing control law pair (V o , k 0 ) and constructs a sequence of cost functional/control law pairs (V i , k i ) in real time, which are stepwise stable and converge to the optimal cost functional/control law pair, for a prescribed nonlinear optimal control problem with unknown input affine state dynamics. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik