Iterative Q -Learning Design for Zero-Sum Games With Evolving Policies
Ding Wang, Yuan Wang, Mingming Zhao, Junfei Qiao · IEEE Transactions on Systems Man and Cybernetics Systems · 2025
This article aims to achieve data-based online evolving control for zero-sum games with unknown dynamics. First of all, the value-iteration-basedQ-learning framework is established. Relevant properties of the iterativeQ-learning framework are analyzed, including the convergence and monotonicity. Then, the stability property is investigated and the online data is employed for off-policy learning. More importantly, two effective algorithms are designed to achieve online evolving control. In one algorithm, the monotonically nondecreasingQ-learning sequence requires the admissible criterion to guarantee the stability with the simpleQ-function initialization. In another algorithm, the monotonically nonincreasingQ-function sequence can ensure the stability without the admissible criterion, but it requires an elaborate initialQ-function. In the end, by including two examples of real physical backgrounds, the excellent performance of online evolving control is exhibited with the given algorithms.