Reinforcement learning in episodic non-stationary markovian environments

Samuel P. M. Choi, Nevin Lianwen Zhang, Dit‐Yan Yeung · Rare & Special e-Zone (The Hong Kong University of Science and Technology) · 2004

Reinforcement learning in non-stationary environments is generally regarded as a very difficult problem. Without any prior knowledge about the environment, this problem can be unsolvable in the worst case. In this paper, we attempt to partially address this grand challenge by formalizing a broad class of non-stationary Markovian environments, of which the state space, action space, transition function, and reward (or cost) function may change over time but with some regularities. We call these environments episodic non-stationary Markovian environments (ENME), which form a fairly common class of non-stationary environments for characterizing many real-world decision problems. We begin with a special subclass of ENMEs called periodic non-stationary Markovian environments (PNME) and then generalize this subclass to more general and realistic forms. Afterwards, we show how the episodic property can be exploited to make the problems solvable by combining conventional reinforcement learning algorithms with the state augmentation method.

Read the paper · More papers on PaperTik