Threshold learning in the improved penalty avoiding rational policy making algorithm

Kazuteru Miyazaki, Ryouhei Kobayashi, Hiroaki Kobayashi · Society of Instrument and Control Engineers of Japan · 2010

The penalty avoiding rational policy making algorithm (PARP) [3] previously improved to save memory and cope with uncertainty, i.e., Improved PARP (IPARP) [4]. The efficiency of IPARP is influenced by threshold of a penalty rule or a penalty basis function γ significantly. In this paper, we propose a technique for learning γ. We show the effectiveness of our proposal using a soccer game task called “Keepaway”[7].

Read the paper · More papers on PaperTik