A Reinforcement Learning Method Based on Adaptive Simulated Annealing

Amir F. Atiya, A.G. Parlos, Lester Ingber · 2006

Reinforcement learning is a hard problem and the majority of the existing algorithms suffer from poor convergence properties for difficult problems. In this paper we propose a new reinforcement learning method that utilizes the power of global optimization methods such as simulated annealing. Specifically, we use a particularly powerful version of simulated annealing called adaptive simulated annealing (ASA) (Ingber, 1989). Towards this end we consider a batch formulation for the reinforcement learning problem, unlike the online formulation almost always used. The advantage of the batch formulation is that it allows state-of-the-art optimization procedures to be employed, and thus can lead to further improvements in algorithmic convergence properties. The proposed algorithm is applied to a decision making test problem, and it is shown to obtain better results than the conventional Q-learning algorithm.

Read the paper · More papers on PaperTik