Learning algorithms for Markov decision processes

Masami Kurano · Journal of Applied Probability · 1987

This study is concerned with finite Markov decision processes whose dynamics and reward structure are unknown but the state is observable exactly. We establish a learning algorithm which yields an optimal policy and construct an adaptive policy which is optimal under the average expected reward criterion.

Read the paper · More papers on PaperTik