Double Deep Q Learning with Gradient Biasing for Mobile Relay Beamforming Networks

Spilios Evmorfos, Konstantinos Diamantaras, Athina P. Petropulu · 2021 55th Asilomar Conference on Signals, Systems, and Computers · 2021

The paper considers the discrete 2D motion control of mobile relays implementing distributed beamforming in a spatiotemporally correlated channel environment. A time-slotted scenario is considered where the relays implement optimal beamforming, while standing still, then estimate the optimal positions for the next slot and move to those selected positions to beamform again. The goal is to maximize the cumulative Signal-To-Interference+Noise Ratio (SINR) at the destination. We employ double deep Q learning to construct the motion policies. The method is completely model free and agnostic of channels statistics. A Fourier feature mapping is applied on the state before passing it to the Q networks, which enables the learning of a richer representation of the Q function in terms of its frequency spectrum. We propose a strategy to bias the neural network gradient updates. In the initial stages of training, our approach induces bias towards easier experiences (experiences that correspond to relatively low loss) from a relay trajectory, while later, it gradually places the bias towards harder examples. This bias transition is controlled by a temperature parameter, that we change through the course of training. The proposed approach provides significant improvement both in reward accumulation and speed of convergence.

Read the paper · More papers on PaperTik