A Reinforcement Learning Method Using Reward Acquisition Efficiency for POMDP Environments
Hirokazu Kawai, Atsushi Ueno, Shoji Tatsumi · Transactions of the Japanese Society for Artificial Intelligence · 2008
Reinforcement Learning (RL) methods are very hopeful because they can learn useful behavior based on rewards from environment by trial and error. This paper tackles more difficult problems than the ones tackled by many ordinary RL methods: RL in POMDP (Partially Observable Markov Decision Process) environments with multiple rewards.