Voronoi model learning for batch mode reinforcement learning

Raphaël Fonteneau, Damien Ernst · ORBi (University of Liège) · 2010

We consider deterministic optimal control problems with continuous state spaces where the information on the system dynamics and the reward function is con-strained to a set of system transitions. Each system transition gathers a state, the action taken while being in this state, the immediate reward observed and the next state reached. In such a context, we propose a new model learning–type reinforce-ment learning (RL) algorithm in batch mode, finite-time and deterministic setting. The algorithm, named Voronoi reinforcement learning (VRL), approximates from a sample of system transitions the system dynamics and the reward function of the optimal control problem using piecewise constant functions on a Voronoi–like partition of the state-action space. 1 Problem statement We consider a discrete-time system whose dynamics over T stages is described by a time-invariant equation xt+1 = f(xt, ut) t = 0, 1,..., T − 1, (1) where for all t ∈ {0,..., T − 1}, the state xt is an element of the bounded normed state space X ⊂ RdX and ut is an element of a finite action space U = a1,..., am with m ∈ N0. x0 ∈ X is the initial state of the system. T ∈ N0 denotes the finite optimization horizon. An instantaneous reward rt = ρ(xt, ut) ∈ R (2) is associated with the action ut ∈ U taken while being in state xt ∈ X. We assume that the initial state of the system x0 ∈ X is fixed. For a given open-loop sequence of actions u = (u0,..., uT−1) ∈ UT, we denote by Ju(x0) the T−stage return of the sequence of actions u when starting from x0, defined as follows: Definition 1.1 (T−stage return) ∀u ∈ UT,∀x0 ∈ X,

Read the paper · More papers on PaperTik