TD Learning of Game Evaluation Functions with Hierarchies of Adaptive Experts

Wiering, B.J.A. Kröse · 1995

Introduction Expert Systems for playing games require an evaluation function which returns the expected payoff from a position given optimal future play of both sides. When the evaluation function for a game is known accurately, it can be used to compare all possible moves in a position after which the move which results in a position with the highest evaluation can be selected. Since neural networks are universal approximators [Cybenko89] they are able to represent the evaluation function. Learning can be supervised (a human expert or a perfect computer program) or by self-play [Tesauro92, Boyan92, Schraudol94] in which the temporal difference method [Sutton88, Dayan94] is used to generate learning samples. Games provide domains where the evaluation function can differ drastically for similar positions (e.g. tic-tac-toe, draughts, chess). To solve the problems of having to represent discontinuities and storing large amounts of incoherent knowledge in one single neural network

Read the paper · More papers on PaperTik