Multi-Agent Deep Reinforcement Learning-Based Resource Allocation for Cognitive Radio Networks
Ruru Mei, Zhugang Wang · IEEE Transactions on Vehicular Technology · 2024
In this paper, we investigate the spectrum access and power control problem in a heterogeneous cellular and vehicular communication network comprising both cellular links and vehicle-to-vehicle (V2V) links. We formulate the spectrum access and power control problem as a fully cooperative multi-agent reinforcement learning task, where each V2V link receives an identical reward at each time step. This strategy caters to the heterogeneous Quality-of-Service requirements, including the high reliability and low latency requirements for safety-centric applications, as well as the high transmission rate requirements for entertainment applications. Consequently, we design a multi-criterion objective reward function that accounts for the heterogeneous traffic demands. Subsequently, we introduce a robust multi-agent proximal policy optimization algorithm with a centralized training and decentralized execution framework. The proposed approach employs global observation to learn policies for each V2V agent during training, while allowing each agent to take an action based on its local observation during distributed execution phase. Furthermore, we consider imperfect spectrum sensing and outdated channel state information (CSI) in the work. Henceforth, a critic network incorporating long short-term memory layers is developed to effectively leverage the global observation, thereby mitigating the impact of outdated CSI and sensing errors. Extensive experiment results demonstrate the superiority of the proposed method comparing to other existing methods in terms of transmission success rate of cellular links, average capacity for entertainment applications of V2V links, and payload delivery rate for safety-centric applications of V2V links. Furthermore, we demonstrate the robustness of the proposed method against variations in V2V payloads and SINR thresholds.