Advantage based value iteration for Markov decision processes with unknown rewards

Pegah Alizadeh, Yann Chevaleyre, François Lévy · 2016

This paper addresses approximating the optimal policy in Markov Decision Process with unknown rewards. The MDP is transformed into a Vector-Valued MDP (VVMDP). We introduce a new interactive algorithm ABVI, whose principle is using value iteration on VVMDPs and querying the user when necessary. This algorithm uses classification method to reduce the number of proposed queries. We integrate value iteration with querying the user to select appropriate backups. In this paper, our goal is to accelerate the value iteration algorithm and to reduce the number of queries.

Read the paper · More papers on PaperTik