Temporal Difference Learning and the Neural MoveMap Heuristic in the Game of Lines of Action

Mark H. M. Winands, Levente Kocsis, JOS W. H. M. UITERWIJK, H.J. van den Herik, Q. Mehdi, N. NGough, M. Cavazza · 2002

This paper investigates to what extent learning methods are beneficial for the Lines of Action tournament program MIA. We focus on two components of the program: (1) the evaluation function and (2) the move ordering. Using temporal difference learning the evaluation function was improved by tuning the weights. We found substantial improvements for three weights. The move ordering was enhanced by the Neural MoveMap (NMM) heuristic, which is based on learning. The two learning techniques improved both the playing quality and the speed of the program. Test results are given. The new evaluation function improved the program with a winning ratio of 1.68. The speed up of the NMM heuristic is 17 percent.

Read the paper · More papers on PaperTik