Control using Q-learning for networked coordination games

Bo Jin, Ming Cao · 2022 13th Asian Control Conference (ASCC) · 2022

We consider a population of individuals playing coordination games on a network, where individuals may differ in their coordination biases and thus the population is heterogeneous. While most of the agents follow the standard asynchronous best-response game-play strategy, a group of agents, called "targeted" agents, are under control and thus can use a different strategy designed to guide the overall population to achieve a pre-specified collective state. The control objective is then to find an optimal policy that specifies which targeted agents should be forced and what strategies they should employ at each time. We formulate this control problem as a Markov decision process and solve it using a Q-learning algorithm. An ergodicity condition turns out to be key to guarantee the convergence to the optimal Q-value. We interpret this ergodicity condition and then further prove that it can be simplified to be checked efficiently when all the individuals’ coordination biases are the same or are all below or all above 1/2.

Read the paper · More papers on PaperTik