Reinforcement Learning Hardware Accelerator using Cache-Based Memoization for Optimized Q-Table Selection
Muhammad Sulthan Mazaya, Eko Mursito Budi, Infall Syafalni, Nana Sutisna, Trio Adiono · 2024
Reinforcement learning (RL) is an approach to building an autonomous agent. RL utilizes a learning mechanism that assigns rewards and punishments for the agent to obtain an optimal policy. Constructing a good-performing rewards and punishments assignment so that the agent has a final converging behavior is a challenging task. Much research has been establishing a guide to build such assignments to address this issue. In this paper, we present a mechanism that allows the RL agent to have a non-converging behavior regardless of its rewards and punishments assignment as a new approach to solving the problem. The mechanism uses Q-table memoization and comparison which introduces overhead. To compensate for the overhead, a hardware accelerator is implemented Kria KV260 Vision AI FPGA using Zynq Ultrascale+ MPSoC soft core processor to accelerate the computing process and tested in a maze reinforcement agent learning setup. It is capable of obtaining latency performance within the range of 21 to 619 clock cycles, which is on average 223.78 times faster than the software implementation. This solution proposed in this research can be used for autonomous systems building and smart navigation systems.