Goal-oriented Semantic Communication in Bandwidth-constrained MARL
Yang Su, Yali Du, Yansha Deng · 2025
Multi-agent reinforcement learning (MARL) with communication has shown remarkable potential in solving complex tasks and promoting cooperation within dynamic environments. However, the communication protocols learned by agents are often difficult to interpret, and important task-related information is not consistently included in communication messages, which can hinder protocol interoperability and negatively impact system performance, especially in environments with limited wireless bandwidth. To address this issue, in this paper, we focus on a collaborative multi-robot navigation scenario with bandwidth-constrained communication and propose a goal-oriented, human-interpretable feature communication protocol. This protocol ensures that features with the richest goal-oriented semantic information are transmitted. We then introduce a novel Multi-Agent Proximal Policy Optimization (MAPPO)-based Bandwidth-Adaptive Dual-Policy (MBDP) Framework, in which a constrained MAPPO-based communication policy selects the most important features for transmission within bandwidth limits, while a MAPPO-based navigation policy generates navigation actions. The two policies are trained alternately. Experimental results demonstrate that our MBDP framework achieves the highest navigation reward compared to other algorithms, efficiently selecting important features to transmit under bandwidth limitations.