Hierarchical MARL with Stackelberg Games

Carmel Fiscko, Haoyu Yin, Bruno Sinopoli · 2025

We consider a multi-agent system under the control of a central planner (CP). We model the system as a hierarchical multi-agent reinforcement learning (MARL) problem in which a state process is influenced by both the CP’s control and the agents’ actions. Each agent learns a local policy to maximize a local value function, which may or may not be aligned with the CP’s control objective. The CP’s goal is therefore to find a global policy such that the agents learn an equilibrium joint policy that maximizes the CP’s value function. We first show that this problem is equivalent to a Stackelberg game. Given a model, this equivalence can be used to derive game-theoretic properties about the desired Stackelberg equilibrium and can be used to solve for optimal policies directly. If the game model is unknown, then we propose a Monte Carlo (MC)-based reinforcement learning (RL) method based on the hierarchical game structure. We demonstrate that under standard RL assumptions, this method can approximate solutions to the desired Stackelberg game. This procedure is validated in simulations on synthetic games resembling social welfare problems.

Read the paper · More papers on PaperTik