Computing bounds for kernel-based policy evaluation in reinforcement learning
Raphaël Fonteneau, Susan Allbritton Murphy, Louis A. Wehenkel, Damien Ernst · Open Repository and Bibliography (University of Liège) · 2010
This technical report proposes an approach for computing bounds on the finite-time return of a policy using kernel-based approximators from a sample of trajectories in a continuous state space and deterministic framework.