3PS: Periodical Partial Policy Sharing for Distributed Cooperative Multi-agent Reinforcement Learning

Ke Zhang, Dandan Zhu, Xiaoning Zhao, Toshiharu Sugawara · Procedia Computer Science · 2025

We propose a framework for multi-agent reinforcement learning called periodic partial policy sharing (3PS) to balance the efficiency of learning and privacy concerns. Because of the negative impacts of instability on team members’ learning, some correct actions are later considered incorrect. To address this problem, current methods employ centralized training and decentralized execution (CTDE) using a shared experience buffer to train consistent policies for individual agents. However, in sensitive applications, agents are not allowed to share their experiences because of privacy concerns. In the proposed 3PS, agents periodically share only the parameters in their policies to mitigate training inefficiencies and prevent privacy breaches from sharing the experience datasets. Depending on the parameters with the associated importance weights of the individual policies, we introduce three sharing methods in 3PS: average 3PS (A-3PS), reward-scalability 3PS (R-3PS), and personalized 3PS (P-3PS). We evaluated these 3PSs using three backbones, QMIX, VDN, and IPPO, for two multi-agent cooperative tasks: SMAC and multi-agent path planning. The results show that 3PS outperforms the baseline methods in terms of the convergence rate, average reward, and win rate, while ofering better data privacy than centralized training. 3PS code is available at https://github.com/ColaZhang22/3PS .

Read the paper · More papers on PaperTik