A class of steering policies under a recurrence condition
Dye-Juan Ma, Armand M. Makowski · 2003
A class of adaptive policies is defined by Markov decision processes (MDPs) under some recurrence conditions. The proposed policy alternates between two stationary policies so as to track adaptively a sample average cost to a desired value. Direct sample path arguments are presented for investigating the convergence of the sample average costs under this adaptive policy. The results have applications to MDPs with a single constraint.>