An OCBA-Based Information Sharing Method for Multi-Agent Reinforcement Learning
Ruicheng Jiang, Qing‐Shan Jia · 2024
Recent technological advancements in renewable energy generation, fuel cells, and energy storage systems have intensified research interest in supply-demand matching within smart grids. This challenge often involves hundreds of thousands of agents making real-time decisions. In the context of multi-agent reinforcement learning (MARL), critical questions arise regarding what information should be shared among agents and how to effectively utilize this shared data. These questions remain largely unresolved, primarily due to the inherent complexity and the variability of the subproblems faced by each agent. We consider this important problem in this work, and make the following major contributions. First, we formulate a framework aimed at maximizing the probability of selecting the optimal action in a given state while operating under a limited sampling budget. Second, we propose an algorithm designed to asymptotically maximize the probability of correct selection (PCS) when all agents are dealing with the same Markov decision process (MDP). Finally, we perform our method on a supply demand matching problem and show its efficiency. We hope this work can bring insights to efficient sampling in MARL in more general situations.