Model-Free Q-Learning for Output Feedback Nash Strategy of Decentralized Nonzero-Sum Games

Qiyan Zhang, Hongxia Wang, Kai Peng, Huanshui Zhang · IEEE Transactions on Systems Man and Cybernetics Systems · 2025

In this article, we present a model-free output feedback (OPFB)Q-learning algorithm to find the optimal Nash equilibrium strategy for the decentralized control problem (DCP) of nonzero-sum games with asymmetric information. The main challenge lies in different historical information available to each controller, namely, the input information is shared while the measurement information is private. To overcome this difficulty, a novel optimal Nash strategy in the input/output form is derived without measurable system states. Then, the OPFBQ-learning iteration algorithm is developed to learn the optimal controllers online only by the knowledge of available input and measurement information, rather than the system dynamics and states. The key is solving the equilibrium equations under asymmetric information, which is achieved by reformulating them into a constrained minimization problem, yielding the numerical solution of the optimal controller pair. The presented idea is new to the best of authors’ knowledge. Numerical examples are shown to illustrate the effectiveness of the proposed algorithm.

Read the paper · More papers on PaperTik