Satisficing Paths and Independent Multiagent Reinforcement Learning in Stochastic Games

Bora Yongacoglu, Gürdal Arslan, Serdar Yüksel · SIAM Journal on Mathematics of Data Science · 2023

Abstract. In multiagent reinforcement learning, independent learners are those that do not observe the actions of other agents in the system. Due to the decentralization of information, it is challenging to design independent learners that drive play to equilibrium. This paper investigates the feasibility of using satisficing dynamics to guide independent learners to approximate equilibrium in stochastic games. For [Formula: see text], an [Formula: see text]-satisficing policy update rule is any rule that instructs the agent to not change its policy when it is [Formula: see text]-best-responding to the policies of the remaining players; [Formula: see text]-satisficing paths are defined to be sequences of joint policies obtained when each agent uses some [Formula: see text]-satisficing policy update rule to select its next policy. We establish structural results on the existence of [Formula: see text]-satisficing paths into [Formula: see text]-equilibrium in both symmetric [Formula: see text]-player games and general stochastic games with two players. We then present an independent learning algorithm for [Formula: see text]-player symmetric games and give high probability guarantees of convergence to [Formula: see text]-equilibrium under self-play. This guarantee is made using symmetry alone, leveraging the previously unexploited structure of [Formula: see text]-satisficing paths.

Read the paper · More papers on PaperTik