Experiments with automatic weight tuning in heuristic evaluation functions

Jónas Tryggvi Jóhannsson · 2007

Tuning weights coefficients in Heuristic Evaluation Function is a non-trivial task, often done by hand by game-playing program developers. The time and effort required to achieve good results has recently driven developers to utilize Reinforcement Learning techniques to tune these weight coefficients automatically. In this masters project, we have developed an general framework for automatic weight tuning. Secondly, we used the framework to conduct a comparisons study between the two main methods of learning from game data, TD(_) and TD-Leaf(_). A world-class chess program called Fruit was modified to enable learning within the framework, and then used to perform the experiments. The results from the study are interesting, as they show that although TD-Leaf(_) is generally believed to be superior and is indeed more robust in respect to different parameter settings, TD(_) can be just as effective when given proper parameter settings.; Þetta verkefnið snerist um að laera betri vigtir fyrir matsfoll (e. Heuristic Evaluation Functions), sem notuð eru i forritum sem spila borðleiki eins og t.d. skak, en til þess var notað skilyrt viðbragð (e. Reinforcement Learning). Verkefnið var tviþaett; annars vegar að bua til almennt forritasafn til að auðvelda ferlið og umstangið i kringum það að laera vigtir fyrir matsfoll. Hinsvegar var forritasafn þetta svo notað til að laera vigtir fyrir mjog sterkt skakforrit sem heitir Fruit-Chess með tveimur mismunandi laerdoms aðferðum sem eru nefndar TD(_) og TD-Leaf(_). Lýsingu a honnun forritasafnsins er að finna i þessari skýrslu eftir að lesendum hefur verið kynnt nauðsynlegt bakgrunnsefni. Niðurstoður tilrauna og samanburð milli aðferðanna er svo að finna i niðurstoðukafla þessarar skýrslu.

Read the paper · More papers on PaperTik