Online Learning Algorithms for Optimal Control and Dynamic Games
Kyriakos G. Vamvoudakis, Frank L. Lewis · 2012
This chapter develops the online adaptive learning algorithms for optimal control and differential dynamic games using measurements along the trajectory. These algorithms are based on actor/critic schemes and involve simultaneous tuning of the actor/critic neural networks and provide online solutions to complex Hamilton-Jacobi equations, along with convergence and Lyapunov stability proofs. The chapter starts by developing an online approximate local smooth solution, based on policy iteration (PI), for the infinite horizon optimal control problem for continuous-time nonlinear systems with known dynamics. It presents an online adaptive algorithm that involves simultaneous tuning of both actor and critic neural networks (i.e., both neural networks are tuned at the same time). One of the major outcomes of the chapter is the online learning algorithm to solve the continuous time multi player nonzero sum games with infinite horizon for linear and nonlinear systems. Controlled Vocabulary Terms optimal control; parallel algorithms