Performance analysis of a new updating rule for TD(λ) learning in feedforward networks for position evaluation in Go game
Horace Wai-kit Chan, Irwin King, John C. S. Lui · 2002
In this paper, a new updating rule for applying temporal difference (TD) learning to multilayer feedforward networks is derived. Networks are trained to evaluate Go board positions by TD(/spl lambda/) learning with different values of /spl lambda/. Performance of each network is estimated by letting it play against other networks. Results show that nonzero /spl lambda/ gives better learning for the network and statistically, larger /spl lambda/ gives better performance.