Model-building adaptive critics for semi-Markov control
Abhijit Gosavi, Susan L. Murray, Jiaqiao Hu, Supriyo Ghosh · 2012
Abstract Adaptive (or actor) critics are a class of reinforcement learning algorithms. Generally, in adaptive critics, one starts with randomized policies and gradually updates the probability of selecting ac-tions until a deterministic policy is obtained. Classi-cally, these algorithms have been studied for Markov decision processes under model-free updates. Algo-rithms that build the model are often more stable and require less training in comparison to their model-free counterparts. We propose a new model-building adaptive critic, which builds the model during the learning, for a discounted-reward semi-Markov de-cision process under some assumptions on the struc-ture of the process. We illustrate the use of our al-gorithm with numerical results on a system with 10 states and a real-world case-study from management science. 1