Proposal and evaluation of the penalty avoiding rational policy making algorithm with penalty level
Kazuteru Miyazaki, Tomomizu Kojima, Hiroaki Kobayashi · 2007
Reinforcement learning (RL) is a kind of machine learning. It aims to adapt an agent to a given environment by utilizing a reward and a penalty. We know the penalty avoiding rational policy making algorithm (PARP) and the penalty avoiding profit sharing (PAPS) as examples of RL systems that are able to suppress a penalty and learn a rational policy. However they cannot treat multiple penalties. In this paper, we extend PARP/PAPS to the environments where there are some kinds of penalties. We propose the penalty avoiding rational policy making algorithm with penalty level (PARPL) that can control how to avoid penalties. We show the effectiveness of PARPLby soccer game simulations.