Online EM algorithm for acquiring evaluation function of game Othello through reinforcement learning

Taku Yoshioka, Shin Ishii · 2003

We previously proposed a method (1998) for acquiring a good evaluation function of the game Othello based on min-max reinforcement learning. In the previous work, we employed a gradient descent method for training the normalized Gaussian network (NGnet) that represents the Othello's evaluation function. However, a lot of games were required to train the NGnet. In this study, we employ the previously proposed online EM algorithm to train the NGnet. The online EM algorithm converges faster than the gradient descent method, and it is suitable for dynamic environments, such as in reinforcement learning tasks. In this article, we introduce a new architecture that is composed of a certain number of NGnets to represent the evaluation function. In this architecture, each of the NGnets is independently trained by the online EM algorithm. Our experiments show that a good evaluation function can be obtained by this new architecture through a smaller number of training games than by the previous scheme.

Read the paper · More papers on PaperTik