Decentralized Q-Learning with Constant Aspirations in Stochastic Games

Bora Yongacoglu, Gürdal Arslan, Serdar Yüksel · 2019

In decentralized stochastic control, coordination among control agents is typically required in order to achieve acceptable system performance. In practice, pertinent information about the system-in the form of the cost function, state transition probabilities, and past actions of other agents-is often unavailable to some or all agents, and this serves as an obstacle to finding optimal control policies. In this paper, a decentralized control problem is modelled as a stochastic game in which (i) the specific game being played is unknown, and (ii) players never observe the actions used by other agents. This information structure exacerbates the already difficult challenge of decentralized policy evaluation in stochastic games. This paper presents a two-timescale reinforcement learning algorithm for stochastic games, in which players engage in decentralized policy evaluation during the finer timescale and update their baseline policies in the coarser timescale. The algorithm presented here comes with provable convergence guarantees and uses only local information, in the form of local cost readings, the history of local actions, and the state information.

Read the paper · More papers on PaperTik