Proposal and Evaluation of the Improved Penalty Avoiding Rational Policy Making Algorithm

Kazuteru Miyazaki, Takuji Namatame, Hiroaki Kobayashi · BiblioBoard Library Catalog (Open Research Library) · 2009

As examples of XoL that use both a reward and a penalty simultaneously, we know PARP and PAPS. The application of PARP to real worlds is difficult since it requires O(MN2) memories where N and M are the number of types of a sensory input and an action. Though PAPS only requires O(MN) memories, it may learn unexpected behavior. In this paper, we have proposed Improved PARP in order to overcome the difficulty by updating a reward and a penalty in each episode and selecting actions depending on the degree of a penalty. We have shown the effectiveness of this approach through a soccer game simulation and its real world experimentation. Improved PARP cannot treat multiple rewards and penalties. In the future, we try to combine Improved PARP with the method in the paper (Miyazaki & Kobayashi, 2004) to treat several types of rewards and penalties at the same time. Furthermore, we will apply Improved PARP to the keepaway task (Stone et.al. 2005) and the other real world applications.

Read the paper · More papers on PaperTik