Resource Allocation in IoT Networks via DQN and PPO Algorithm

Shrishti Garg, Shruti Gupta, Vivek Prakash Srivastava · 2025

Since the last couple of years has witnessed a steadily increasing demand for IoT-based services, their corresponding cloud infrastructures must distribute and allocate resources in efficient and dynamic manners. Workloads from the Internet of Things devices are diverse as well as changeable, so traditional resource allocation strategies are not suitable. In this work, the problem of dynamic resource allocation in cloud environments for datasets on the Internet of Things is explored and applied with RL techniques, such as DQN and PPO. The dynamic state of cloud resources, such as CPU, memory, and bandwidth usage, is introduced. By considering resource allocation as a Markov Decision Process (MDP), an investigation was made regarding how the RL algorithm can optimize the use of resources while adhering to SLAs. DQN value-based reinforcement learning technique is used to evaluate the best allocation rules by estimating Q-values of state-action pairs. On the other hand, PPO is a policy-gradient method used to optimize the policy directly and provide stability and robustness in continuous action spaces. In this evaluation, an Internet of Things-based dataset was used to evaluate both algorithms regarding dynamic modification of allocation resources few SLA breaches, and improved resource utilization efficiency. A comparative result shows that PPO does a better job compared to DQN in scaling any kind of resource continuously with higher rewards and faster convergence compared to the former approach. This work demonstrates the potential for RL-based techniques for cloud resource management within the Internet of Things ecosystems.

Read the paper · More papers on PaperTik