Addressing the Policy-bias of Q-learning by Repeating Updates

Sherief Abdallah, Michael Kaisers, Ito, Tsuyoshi, Jonker, Catholijn, Gini, M., Shehory, Onn · 2013

Q-learning is a very popular reinforcement learning algorithm be-ing proven to converge to optimal policies in Markov decision pro-cesses. However, Q-learning shows artifacts in non-stationary en-vironments, e.g., the probability of playing the optimal action may decrease if Q-values deviate significantly from the true values, a sit-uation that may arise in the initial phase as well as after changes in the environment.These artifacts were resolved in literature by the variant Frequency Adjusted Q-learning (FAQL). However, FAQL also suffered from practical concerns that limited the policy sub-space for which the behavior was improved. Here, we introduce the Repeated Update Q-learning (RUQL), a variant of Q-learning that resolves the undesirable artifacts of Q-learning without the practi-cal concerns of FAQL. We show (both theoretically and experimen-tally) the similarities and differences between RUQL and FAQL (the closest state-of-the-art). Experimental results verify the theo-retical insights and show how RUQL outperforms FAQL and QL in non-stationary environments.

Read the paper · More papers on PaperTik