Rollout-based Shapley Values for Explainable Cooperative Multi-Agent Reinforcement Learning
Franco Ruggeri, William Emanuelsson, Ahmad Terra, Rafia Inam, Karl Henrik Johansson · 2024
Credit assignment in cooperative Multi-Agent Reinforcement Learning (MARL) focuses on quantifying individual agent contributions toward achieving a shared objective. One widely adopted approach to compute these contributions is through the application of Shapley Values, a concept derived from game theory. Previous research in Explainable Reinforcement Learning (XRL) successfully computed Global Shapley Values (GSVs), albeit neglecting local explanations in specific situations. In contrast, another approach concentrated on learning Local Shapley Values (LSVs) during training, prioritizing sample efficiency over explainability. In this paper, we extend an existing method to generate local and global explanations in a model-agnostic manner, bridging the gap between these two approaches. We apply our proposed algorithm to two cooperative tasks: a predator-prey environment and an antenna tilt optimization problem in cellular networks. Our findings reveal that the LSVs offer valuable insights into the agents’ behavior with a finer time-frame granularity, while their aggregation in GSVs enhances trust by potentially identifying suboptimality. Importantly, our approach surpasses the existing state-of-the-art methods in estimating LSVs, enhancing the accuracy of assessing individual agent contributions. This work represents a significant advancement in the field of XRL and provides a powerful tool for gaining deeper insights into agents’ behavior in cooperative MARL systems.