Solving Go on a 3x3 Board Using Temporal-Difierence Learning

Choon Ngai Tay · 2004

Go is a fascinating board game that originated from China more than 3000 years ago. Even after decades of research and lots of programming time spent on Go, the best computer Go is still playing at a modest level. On the other hand, other computer board games, such as chess, checkers, Othello and backgammon have attained grandmaster level. There are two reasons associated with the computer Go’s unsuccessful performance. The first reason is the enormous search space for Go, which is approximately 10 states, and the second and main reason is due to the absence of a good evaluation function that is able to accurately describes the Go state. Many AI techniques that were used and worked on other board games were attempted on Go but without equivalent success. Following the successful application of TD-learning on backgammon, researchers began studies applying TD-learning to Go. Their learning programs with no prior knowledge encoded managed to acquire ‘genuine’ knowledge and managed to achieve the level of a low level playing Go commercial program. This motivated our investigation into the applications of TD-learning on Go. As Go on small boards, such as 2 x 2, 3 x 3, and 4 x 4 were solved using game tree search, we attempted to solve a 3 x 3 board using another approach — TD-learning. We developed learning agents using TD(0), the simplest form of TD(λ), and TD-directed(0) each with a lookup table, and trained these against some selfdeveloped training agents. After 60000 training games, we observed that all agents were learning and some even managed to secure a 100% win ratio against a non-prefect test agent. We discuss further enhancements that can be made to improve the performance of the learning Go program for solving small boards, and possible extensions to solve bigger boards.

Read the paper · More papers on PaperTik