Distributed MARL for Scheduling in Conflict Graphs
Yiming Zhang, Dongning Guo · 2023
This paper addresses a link scheduling problem in networks represented by conflict graphs using a distributed learning approach. Each agent in a network controls a single link and has access only to its own state and the states of links in its neighborhood. The goal is to minimize the average packet delay in the multi-agent network. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). The proposed solution adopts a centralized training and distributed execution paradigm and leverages an on-policy reinforcement learning algorithm. Specifically, the paper employs the multi-agent proximal policy optimization (MAPPO) algorithm with judiciously designed recurrent structures in the neural network. The proposed solution is shown to outperform some widely used schedulers in terms of throughput and delays through simulations.