Non-Randomized Markov and Semi-Markov Strategies in Dynamic Programming
E. A. Fainberg · Theory of Probability and Its Applications · 1982
In a non-homogeneous controllable Markov model with a total reward criterion, discrete time, infinite horizon and Borel spaces of states and controls, let a certain strategy 7r and an initial measure /x be given. In the paper the following two statements are proved: