Bayes Adaptive Reinforcement Learning versus Off-line Prior-based Policy Search: an Empirical Comparison

Michaël Castronovo, Damien Ernst, Raphaël Fonteneau · 2014

This paper addresses the problem of decision making in unknown nite Markov Decision Processes (MDPs). The uncertainty about the MDPs is modelled, using a prior distribution over a set of candidate MDPs. The performance criterion is the expected sum of discounted rewards collected over an innite length trajectory. Time constraints are dened

Read the paper · More papers on PaperTik