Fast Reinforcement Learning using multiple models
Kumpati S. Narendra, Yu Wang, Snehasis Mukhopadhay · 2016
Reinforcement Learning aims to find the optimal decision in uncertain environments on the basis of qualitative and noisy on-line performance feedback provided by the environments. During the past four decades, learning theory has grown into a vast field in which a very large number of problems have been studied. One of the primary limitations of reinforcement schemes, acknowledged by workers in the field, is their slow speed of convergence. The principal objective of this paper is to present a new approach, based on the use of multiple models (or estimates), that may alleviate this problem and increase the speed of response. In adaptive control theory, multiple model based methods have been proposed over the past two decades, which improve substantially the performance of the system. The authors undertook to apply similar concepts in reinforcement learning as well, and this paper represents the first effort in this direction. Simple situations of learning in feed-forward networks are considered in the paper, and compared to two different schemes. It is shown that convergence speeds that are more than an order of magnitude faster than those of the first scheme, can be achieved in some cases. While the second scheme is comparable to the new approach in many situations, it is seen to exhibit undesirable behaviour in others, where the new approach is more robust. The latter is currently being extended incrementally and systematically to more complex problems that have been discussed in the literature. The ultimate aim of the authors is to apply this approach to learning in discrete and continuous state dynamic environments.