On Sampled Reinforcement Learning in Wireless Networks: Exploitation of Policy Structures

Libin Liu, Urbashi Mitra · IEEE Transactions on Communications · 2020

Reinforcement learning is a classical tool to solve network control or policy optimization problems in unknown environments. In order to learn the optimal policy correctly, the classical Q-learning algorithm requires sufficient visits to all state-action pairs, resulting in the need for a large number of observations in the presence of a large state-action space. Nevertheless, complexity reduction can be achieved by exploiting the particular structure of the optimal policy. A sampled reinforcement learning algorithm is proposed, where the optimal policy is estimated only for a subset of states; a machine learning technique, as well as a graph signal processing approach, are applied for policy interpolation for unvisited states. A policy refinement algorithm is further proposed to improve the performance of policy interpolation. Performance analysis and bounds are also provided for the proposed policy sampling and interpolation algorithms. Numerical experiments on a single link wireless network with a large state space show that the sample Q-learning algorithm with policy interpolation achieves a much faster runtime with negligible performance loss compared to classical Q-learning.

Read the paper · More papers on PaperTik