Improving Decision-Making Policy for Cognitive Radio Applications Using Reinforcement Learning
Rajat Singh, Jayant Kumar Rai, Pinku Ranjan, Rakesh Chowdhury · 2024
Motivated by cognitive radios, there has been a recent increase of interest in stochastic multi-player multi-armed bandits. In this context of cognitive radio’s, autonomous players concurrently engage in arm or channel pulls, individually opti-mizing rewards. Complexity amplifies with potential collisions, wherein multiple players simultaneously select a common arm, resulting in zero collective reward. Our work centers on the Multiplayer Multi-Armed Bandit (MMAB) problem, involving M decision makers collaborating to maximize cumulative reward in cognitive radio application. Collision prompts players to adapt. We introduce RobustMMAB, a decentralized algorithm aiming to achieve regret akin to an optimal centralized algorithm while increasing resilience against selfish nodes.