MIXING OF MARKOV PROCESSES*
Edward A. Silver, John B. Moore · Decision Sciences · 1976
ABSTRACT This paper considers a Markov process in which there are rewards associated with state transitions. Two or more actions are available to the decision maker, and the transition probabilitities and the rewards are influenced by the actions selected. A key assumption is that control cannot be a function of the current state of the process. In other words, the decision maker is assumed to have no knowledge of the current state of the process (or to ignore such knowledge if it exists). Within this context the problem of determining the best randomized policy, the best in the sense of maximizing the expected reward per transition in the steady state, is discussed. Analytic results are developed for the simplest case of a two state process with two possible actions.