Exploiting policy structure for solving MDPs with large state space
Libin Liu, Arpan Chattopadhyay, Urbashi Mitra · 2018
Markov decision processes provide good models for many systems, including wireless communication networks. The goal herein is to develop optimal control policies for wireless networks. While classical methods such as value iteration and policy iteration have been employed to determine optimal policies in a moderate complexity manner; they still suffer from complexity challenges for very large scale networks. Previously, subspace approximation has been employed to find optimal controllers in reduced dimensions. Herein, an alternative approach is considered wherein the properties of the policy structure are exploited to determine solutions in a reduced dimension. The numerical results show that this new approach achieves a faster convergence rate with a negligible loss of performance.