Multi-criteria Reinforcement Learning

Zoltán Gábor, Zsolt Kalmár, Csaba Szepesvári · 1998

We consider multi-criteria sequential decision making problems where the vector-valued evaluations are compared by a fixed total ordering of the vectors. Conditions for the optimality of stationary policies and the Bellman optimality equation are given for a special, but important class of problems, when the evaluation of policies can be computed componentwise. The analysis requires special care as the topology introduced by pointwise convergence and the order-topology introduced by the preference order are in general incompatible several. Reinforcement learning algorithms are then proposed and analyzed. Preliminary computer experiments confirm the validity of the derived algorithms. These type of multi-criteria problems are most useful when there are several optimal solutions to a problem and one wants to choose the one among these which is optimal according to another fixed criterion. Possible application in robotics and repeated games are outlined. 1 Introduction Scalar-valued rein...

Read the paper · More papers on PaperTik