DQN Algorithm Design for Fast Efficient Shortest Path System
A Sumarudin, Nana Sutisna, Infall Syafalni, Bambang Riyanto Trilaksono, Trio Adiono · 2023
Reinforcement Learning is an algorithm with decision-making based on the Markov Decision Process (MDP). MDP uses actions that are assessed by rewards to achieve the goal depending on the environment. An action is taken by a policy. In this work, we consider a maze-based shortest path problem. The maze problem is solved based on the value of Q from the chosen action. The ordinary Q-learning algorithm uses these value-based states in two modes; explore and exploit. However, the continuous states require initialization. The problem appears in very large states. Thus, in this work, deep learning in the form of neural-network is embedded to determine the Q-value called by Deep Q-Network (DQN). The neural network architecture has a different number of layers and neurons for each implementation. Moreover, the implementation of agents on edges has limited resources. Thus, it requires an optimal architecture to achieve the goal. This study uses an environment maze with a size of 20 × 20 with four actions; up, down, left, and right to reach the destination state or goal. In this paper, we study how to get the optimum processing elements (PE) to be implemented in limited resources of edge computing. Experimental results show that the minimum number of layers of three layers with 14 hidden neurons, and 200 deep replay memory from DQN architecture have been able to complete the shortest path on a maze with a size of 20 × 20. The finding effectively minimizes the amount of computing and potential for edge computing applications.