Optimality equations in undiscounted Markov decision processes
Martin L. Puterman · 2003
We explore properties of the average and bias optimality equations in unichain Markov decision processes. We show that in unichain models, these equations have the same form, so that theory for gain optimality carries over to bias optimality.