An expected experience replay based Q-learning algorithm with random state transition
Feng S. Zhang, Hui QIAN, Chun-Ru Dong, Qiang Hua · JOURNAL OF SHENZHEN UNIVERSITY SCIENCE AND ENGINEERING · 2020
The experience replay method in reinforcement learning algorithms reduces the correlation between state sequences by sampling randomly and increases the efficiency of data utilization. However, presently it can only be used in the deterministic environment. In order to use the experience replay efficiently in a dynamic random environment and keep the original state transition distribution unchanged, we propose a tree-based experience storage structure to store the state transition probability in the process of exploration and provide an expected experience replay based Q-learning algorithm which realizes an unbiased estimation of transition distribution. The main advantage of proposed algorithm lies in that it can keep the transition distribution unchanged without increasing the algorithm complexity. Additionally, it eliminates the overestimation of Q value in an efficient way. Experimental results in the classical random walking problem of robot verify that the proposed algorithm improves the convergence speed by about 50%.