Dedicated TD-learning for Stronger Gameplay: applications to Go

R Ekker, Ecd van der Werf, Lambert Schomaker · University of Groningen research database (University of Groningen / Centre for Information Technology) · 2004

This paper presents a study of several dedicated Temporal-Difference (TD) reinforcement learning algorithms for deterministic zero-sum games of perfect information such as the game of Go. The algorithms include TD(mu) by Beal (2002), which separates good play from bad play, TD-leaf(lambda) and TD-directed(lambda) by Baxter et al. (1998), which exploit game tree searching, as well as Baird's residual algorithm (1995) for preventing instability during training. We show that dedicated TD learning algorithms provide faster training and the acquisition of more 'genuine' knowledge of the game resulting in significantly higher playing strength than players trained by standard TD.

Read the paper · More papers on PaperTik