Stable Fitted Reinforcement Learning
Geoffrey J. Gordon · 1995
We describe the reinforcement learning problem, motivate algorithms which seek an approximation to the Q function, and present new convergence results for two such algorithms. 1 INTRODUCTION AND BACKGROUND Imagine an agent acting in some environment. At time t, the environment is in some state x t chosen from a finite set of states. The agent perceives x t , and is allowed to choose an action a t from some finite set of actions. The environment then changes state, so that at time (t + 1) it is in a new state x t+1 chosen from a probability distribution which depends only on x t and a t . Meanwhile, the agent experiences a real-valued cost c t , chosen from a distribution which also depends only on x t and a t and which has finite mean and variance. Such an environment is called a Markov decision process, or MDP. The reinforcement learning problem is to control an MDP to minimize the expected discounted cost P t fl t c t for some discount factor fl 2 [0; 1]. Define the function Q ...