A Quantum Temporal Difference Learning Method Based on Quantum World Model
Peigen Zeng, Ying He, Fei Richard Yu, Jianbo Du · 2024
Based on quantum parallelism theory and quantum phenomena such as superposition and entanglement, quantum reinforcement learning (QRL) has the potential to surpass classical reinforcement learning (RL). Although some excellent works have been done on QRL, existing RL methods encounter chanllenges when performing in environments with sparse rewards. In this paper, we provide a new perspective on conducting temporal difference (TD) learning in quantum computing, which can eliminate redundant exploration steps compared to classical methods. Specifically, we first use environment information to construct a world model with quantum circuits, enabling it to interact in a quantum way. Then, we perform the learning process by using the quantum world model and Grover’s algorithm to query backwards how to reach the recorded states with high TD-errors. Simulation results show that our proposed method has superior performance compared to classical RL algorithms.