Online Solution of +onlinear TwoPlayer ZeroSum Ga mes Using Synchronous Policy Iteration

Kyriakos G. Vamvoudakis, Frank L. Lewis · 2010

In this paper we present an online gaming algorithm based on policy iteration to solve the continuoustime (CT) twoplayer zerosum game with infinite horizon cost for nonlinear systems with known dynamics. That is, the algorithm learns online in realtime the solution to the game design HJI equation. This method finds in realtime suitable approximations of the optimal value, and the saddle point control policy and disturbance policy, while also guaranteeing closedloop stability. The adaptive algorithm is i mplemented as an actor/critic structure which involves simultaneous continuoustime adaptation of critic, control actor , and disturbance neural networks. We call this online gaming algorithm 'synchronous' zerosum game policy iterat ion. A persistence of excitation condition is shown to guarantee convergence of the critic to the actual optimal value function. +ovel tuning algorithms are given for critic, actor and disturbance networks. The convergence to the optimal saddle point solution is proven, and stability of the system is also guaranteed. Simulation examples show the effectiveness of the

Read the paper · More papers on PaperTik