Proximal Policy Optimization in Uncoordinated and Distributed Multi-Agent Resource Allocation in Cognitive Radio Network
Ankita Vijay Tondwalkar, Andres Kwasinski · 2025
This paper studies the learning performance of the Proximal Policy Optimization (PPO) algorithm in the challenging use case of uncoordinated and distributed multi-agent (UDMA) resource allocation in a cognitive radio network (CRN) via an underlay dynamic spectrum access (DSA) paradigm. The challenge is to establish whether PPO will converge faster to the best performance policies than UDMA-DQL in the non-stationary environment. The simulation results show that the UDMA-PPO achieves faster learning performance than UDMA table-based Q-learning and UDMA deep Q-learning (DQL) and is robust to the non-stationary environment. The UDMA-PPO achieves better performance with about 70 % less training steps compared to the UDMA-Table and achieves similar performance with 25% less training steps than the UDMA-DQL algorithm.