Twenty-first international conference on Machine learning - ICML '04

2004

We consider learning in a Markov decision process where we are not explicitly given a reward function, but where instead we can observe an expert demonstrating the task that we want to learn to perform. This setting is useful in applications (such as ...

Read the paper · More papers on PaperTik