Improved Reinforcement Learning for Resource Allocation in Multi-User Multiple-Input Multiple-Output Networks
Pavan Srikanth SubbaRaju Patchamatla, Haydeer MohamadAbbas, Veeranna Kotagi, G. Durgadevi, Arul Leo Felix L · 2025
Nowadays, the increasing demand for high-speed and low latency wireless communication has driven the development of multi-user Multiple-Input Multiple-Output (MIMO) networks. The existing Q-learning has been used to optimize the resource allocation policies through interactions with the environment. However, it has led to significant challenges such as high-dimensional state and action spaces. Hence, this research proposes Improved Reinforcement Learning (IRL) based on Actor-Critic framework to allocate resources in multi-user MIMO networks by adapting to changing network conditions. Initially, the channel is designed using the standardized 3GPP 3D-Uma (Urban Macro) channel model and the ray-tracing data is used to ensure accuracy. Here, the proposed framework consists of pointer network as the policy network and a value network. The pointer network takes Channel State Information (CSI) as input to generate a sequence of user indices directly and this sequential decision-making capability allows context aware user scheduling. Then, the value network evaluates the performance of selected users and finally, the framework is trained using policy gradient as well as value function estimation. The proposed IRL achieved better results in terms of spectral efficiency for user number 20 (3.52bps/Hz) as well as 30 (3.16bps/Hz) and running time for user number 20 (8.7s) and 30 (7.9s) respectively when compared to existing RL.