Multiagent Learning by Iterative Refinement of Game Models

Yongzhao Wang · Deep Blue (University of Michigan) · 2023

The methodology of Empirical Game-Theoretic Analysis (EGTA) offers a comprehensive collection of techniques for game reasoning with models based on simulation data. For multiagent systems not amenable to analytic solution, EGTA provides a simulation-based alternative, where a game model with a selected set of strategies is evaluated, addressing the most important strategic considerations. The challenge of efficiently assembling a suitable collection of strategies for a game model in EGTA is called the strategy exploration problem. The clearest formulation of strategy exploration in EGTA is within an iterative process, in which a game model is iteratively refined through the alternation of the creation of new strategies and the assessment and analysis of the current game model. In particular, the Policy Space Response Oracles (PSRO) algorithm provides a flexible framework for strategy exploration, with new strategies generated each iteration through a best response to a target other-players profile using reinforcement learning (RL). The component responsible for determining the target profile is called a meta-strategy solver (MSS), which takes an empirical game model as input and "solves'' it to produce the target. I actively investigate three main research aspects of strategy exploration under the PSRO framework: evaluating strategy exploration, controlling strategy exploration, and extension of strategy exploration to mean field games (MFGs). First, I investigate some of the methodological considerations in evaluating intermediate game models generated through strategy exploration, proposing and justifying new evaluation methods based on examples and experimental observations. In particular, I emphasize the fact that empirical games create a space of strategies and evaluation should reflect how well it covers the strategically relevant space. Based on this fact, I propose a new evaluation scheme that measures the strategic coverage of an empirical game. Second, I investigate how to control strategy exploration to build a game model that involves desired solutions of the full game with minimum computational costs. Specifically, I investigate controlling strategy exploration by setting MSSs. I introduce a novel MSS for PSRO, called regularized replicator dynamics (RRD), which prevents overfitting by terminating replicator dynamics before it reaches an exact Nash equilibrium (NE). I demonstrate the effectiveness of RRD on identifying strategically important strategies and accelerating strategy exploration in games with large strategy spaces. I investigate an alternative means for the controlling: setting the response objective (RO) employed in deriving a strategy for a given target profile. My motivation is that different ROs may steer strategy exploration toward solutions with various desired properties. I perform a study in the domain of sequential bargaining games, comparing the standard RO based on own payoff with others based on social welfare. I find that an RO encoded with Nash product can lead to identifying equilibrium outcomes with significantly higher social welfare than the standard objective. Third, I extend the iterative EGTA framework to MFGs. I first prove the existence of NE in the empirical MFG, which then serves as the MSS in the framework. I introduce a game model learning approach, which is essentially a form of regression of the utility function based on utility data collected from previous EGTA iterations. A learned utility function can generalize across mean fields and thus completing the definition of a game model. I combine the iterative EGTA framework with game model learning and provide an effective and sample efficient EGTA framework for MFGs.

Read the paper · More papers on PaperTik