Research on Resource Allocation Model of Reinforcement Learning in Wireless Communication Networks

Kuang Hong, Qian Ke · 2025

This paper proposes a combination of Proximal Policy Optimization (PPO), Simulated Annealing (SA) algorithm and Time-sensitive Network (Time-Sensitive Network). The reinforcement learning method of TSN - Proximal Annealing Sensitive Network (PASN), is used for resource allocation in wireless communication networks. This method aims to improve the efficiency and fairness of resource allocation while reducing energy consumption. Firstly, the efficient policy update mechanism of the PPO algorithm is utilized to rapidly optimize the resource allocation strategy. Secondly, the global search ability of the SA algorithm is introduced to avoid getting trapped in local optimal solutions and ensure the global optimality of resource allocation. Finally, combined with the time sensitivity modeling of the network state by TSN, the resource allocation is dynamically adjusted to adapt to the real-time changes of network traffic. The experimental results show that, compared with traditional methods, PASN can allocate resources more efficiently, reduce energy consumption, and improve the overall performance of the network. This method effectively enhances the flexibility of the resource allocation strategy in a dynamic network environment and provides a new technical path for the intelligent scheduling of wireless communication networks.

Read the paper · More papers on PaperTik