Q_Learning based on active backup and memory mechanism
Yang Liu, Maozu Guo, Hongxun Yao · 2004
Exploration is used in Q/spl I.bar/learning because the agent would be caught in locally optimal policies due to blind exploitation. However excessive exploration would degrade the performance of Q/spl I.bar/learning and it is difficult to meet the trade-off between exploration and exploitation. The active backup is introduced into Q/spl I.bar/learning and the corresponding algorithm AB/spl I.bar/Q/spl I.bar/learning based on Dijkstra backup in dynamic programming is proposed. Then, the memory mechanism based MEAB/spl I.bar/Q/spl I.bar/Iearning algorithm is given for the agent to learn in completely unknown environment. The experimental results show that these two algorithms not only converge more quickly, but also solve the problem of local optimization.