Labeling Q-learning for non-Markovian environments
Hae Yeon Lee, Hiroyuki Kamaya, K. Abe · 2003
The most widely used reinforcement learning (RL) algorithms, such as Q-learning and TD (/spl lambda/) are limited to Markovian environments. Recent research on reinforcement learning algorithms has concentrated on partially observable Markov decision process (POMDP). The only way to overcome partial observability is to use memory to estimate state. In this paper, we present a new memory architecture of RL algorithms to solve certain type of POMDPs. Our algorithm, which we call labeling Q-learning (LQ-learning), is applied to test problems of simple mazes taken from recent literature. The results demonstrate LQ-learning's ability to work well in near optimal manner.