Relaxation schemes for min max generalization in deterministic batch mode reinforcement learning

Raphaël Fonteneau, Damien Ernst, Bernard Boigelot, Quentin Louveaux · Open Repository and Bibliography (University of Liège) · 2011

We study the minmax optimization problem introduced in [6] for computing policies for batch mode reinforcement learning in a deterministic setting. This problem is NP-hard. We focus on the two-stage case for which we provide two relaxation schemes. The first relaxation scheme works by dropping some con-straints in order to obtain a problem that is solvable in polynomial time. The second relaxation scheme, based on a Lagrangian relaxation where all constraints are dualized, leads to a conic quadratic programming problem. Both relaxation schemes are shown to provide better results than those given in [6]. 1

Read the paper · More papers on PaperTik