Lookahead planning and co-evolution in recurrent neural networks
Yuji Sato, S. Hatano, Hisaaki Hatano, T. Furuya · 2002
This paper describes an investigation into the effectiveness of a lookahead model based upon a recurrent neural network. An action network and an internal model of the environment are incorporated into the recurrent neural network, and lookahead planning is performed while configuring the action network through learning in the internal model. In addition, the lookahead planning results are used in the learning process of the action network. In other words, the two networks undergo co-evolution. A genetic algorithm is applied to the construction of the neural network and to the learning of weights. The effectiveness of this model is evaluated by applying it to the game of "tic-tac-toe," and the following conclusions are obtained: (i) By performing lookahead planning using an internal model of the environment, it is possible to reduce the number of trial cycles required for learning from a real environment. (ii) The internal model should only be switched in when the learning process has progressed beyond a certain level. (iii) It is possibly more effective to perform learning in the internal model by learning algorithms than by memorizing input-output correspondences.>