Seeking equilibrium for linear-quadratic two-player Stackelberg game: a Q-learning approach
曼 李, 家虎 秦, 龙 王 · Scientia Sinica Informationis · 2021
In recent years, Stackelberg game has contributed a lot to security control of cyber-physical systems and to energy management in smart grids.The existing methods for seeking Stackelberg equilibrium rely heavily upon complete information of the system dynamics; however, exact system dynamics is difficult to get in real applications, which restricts the applications of the theoretical research results to some extent.In view of this, this paper proposes to seek the equilibrium for Stackelberg game in a model-free way.Specifically, we investigate the linear-quadratic two-player Stackelberg game, in which the game state is evolved along with a linear system and the cost functions are quadratic.The two players in this game are called leader and follower, where the leader makes its decision preferentially with consideration of the reaction functions of the follower, while the follower reacts optimally to the leader's strategy.Due to the consideration of linear state dynamics and quadratic cost functions, as well as the fact that the leader takes actions prior to the follower, the decision-making problem for the leader and the follower can be formulated as a two-level linear-quadratic optimal control problem.According to the principle “from the follower to the leader”, this paper derives a pair of optimal control strategies through dynamic programming.The resulting strategies are shown exactly to be the Stackelberg equilibria, but they depend on the information of system dynamics.Then a new actor-critic based Q-learning algorithm, which could approximate the resulting equilibrium strategies without any information of system dynamics, is proposed.It is shown that under the proposed Q-learning algorithm, the system state as well as the approximation errors of the parameters for actor and critic neural networks are uniformly ultimately bounded.The simulation results show that the control strategies obtained from the proposed Q-learning algorithm could make the system state stable, and the cost functions under the estimated control strategies have only a small deviation from the optimal ones.