Data-Driven Q-Learning in Dynamic Environment

Guoyin Wang · Journal of Southwest Jiaotong University · 2009

It is difficult for reinforcement learning to balance between the exploration of untested actions and the exploitation of known optimum actions in dynamic environment.To address this problem,a data-driven Q-learning algorithm was proposed.In this algorithm,the information system of behavior is constructed for each agent.Then the trigger mechanism of environment is build by the uncertainty of knowledge in the information system of behavior to trace the environmental change.The dynamic information of the environment is used to exploit new environment by the trigger mechanism to achieve the balance between the exploration of untested actions and the exploitation of know optimum actions.The proposed algorithm was applied to grid-world navigation tasks.The simulation results show that compared with the Q-learning,simulated annealing Q-learning(SAQ) and recency-based exploration(RBE) Q-learning algorithms,the proposed algorithm has a high learning efficiency.

Read the paper · More papers on PaperTik