Sample path-based policy-only learning by actor neural networks
Eiji Mizutani · 2003
This paper highlights a sample-path-based policy-only learning algorithm for regenerative (stochastic) processes proposed by Marbach and Tsitsiklis (1998). The algorithm attempts to optimize a randomized, parameterized policy according to the average cost criteria in conjunction with the so-called infinitesimal perturbation analysis gradient estimation technique. We present our numerical studies, demonstrating this learning algorithm using small-scale problems; in particular, the parametrized policy-only agent is a neural network function approximator in the spirit of neuro-dynamic programming (or reinforcement learning).