PAC learning for markov decision processes and dynamic games

Rahul Kumar Jain, Pravin P. Varaiya · 2004

We extend the probably approximately correct (PAC) model of learning to Markov decision processes (MDPs) and dynamic games. We obtain simulation-based uniform sample complexity bounds for value function estimates of discounted reward MDPs. We also obtain uniform sample complexity results for Markov games with a finite number of players.

Read the paper · More papers on PaperTik