A Learning Algorithm for Risk-Sensitive Cost

Arnab Basu, Tirthankar Bhattacharyya, Vivek S. Borkar · Mathematics of Operations Research · 2008

A linear function approximation-based reinforcement learning algorithm is proposed for Markov decision processes with infinite horizon risk-sensitive cost. Its convergence is proved using the “o.d.e. method” for stochastic approximation. The scheme is also extended to continuous state space processes.

Read the paper · More papers on PaperTik