Reinforcement Learning based Path Selection in Multi-hop Device to Device Communication
Rashmita Routray, Tapas Kumar Patra · 2024
This paper investigates the computation time required to establish a multi-hop path between two devices using a reinforcement learning algorithm. The goal is to relieve the base station from traffic burdens and increase network efficiency. Path selection plays a crucial role in Device-to-Device (D2D) communication to reduce time delays in data transmission. Here, we attempt to establish a path by creating a multi-hop routing path. The device in the multi-hop path will act as a relay and establish D2D links with both the source and destination nodes. In this paper, the Dijkstra algorithm and reinforcement learning are used to find an optimum path for multi-hop communication. A hexagonal cell scenario is considered here by taking 25 numbers of cellular users, and the computation time for the number of communication requests is calculated using both algorithms. The reinforcement learning algorithm has a minimum computation time of 25 microseconds as compared to the Dijkstra algorithm. The computation time guarantees that the reinforcement learning algorithm minimizes time delay.