A Reinforcement Learning Algorithm in Partially Observable Environments Using Short-Term Memory
Nobuo Suematsu, Akira Hayashi · 1998
We describe a Reinforcement Learning algorithm for partially observable environments using short-term memory, which we call BLHT. Since BLHT learns a stochastic model based on Bayesian Learning, the overfitting problem is reasonably solved. Moreover, BLHT has an efficient implementation. This paper shows that the model learned by BLHT converges to one which provides the most accurate predictions of percepts and rewards, given short-term memory. 1 INTRODUCTION Research on Reinforcement Learning (RL) probHTM POMDP model-free d b a c e World Policy Figure 1: Three approaches lem for partially observable environments is gaining more attention recently. This is mainly because the assumption that perfect and complete perception of the state of the environment is available for the learning agent, which many previous RL algorithms require, is not valid for many realistic environments. One of the approaches to the problem is the model-free approach (Singh et al. 1995; Jaakkola et al. 1995) ...