Probability Redistribution using Time Hopping for Reinforcement Learning
Petar S. Kormushev, Fangyan Dong, Kaoru Hirota · Spiral (Imperial College London) · 2009
—A method for using the Time Hopping technique as a tool for probability redistribution is proposed. Applied to reinforcement learning in a simulation, it is able to re-shape the state probability distribution of the underlying Markov decision process as desired. This is achieved by modifying the target selection strategy of Time Hopping appropriately. Experiments with a robot maze reinforcement learning problem show that the method improves the exploration efficiency by re-shaping the state probability distribution to an almost uniform distribution.