An Actor/Critic Algorithm that is Equivalent to Q-Learning

Robert H. Crites, Andrew G. Barto · 1994

We prove the convergence of an actor/critic algorithm that is equivalent to Q-learning by construction. Its equivalence is achieved by encoding Q-values within the policy and value function of the actor and critic. The resultant actor/critic algorithm is novel in two ways: it updates the critic only when the most probable action is executed from any given state, and it rewards the actor using criteria that depend on the relative probability of the action that was executed. 1 INTRODUCTION In actor/critic learning systems, the actor implements a stochastic policy that maps states to action probability vectors, and the critic attempts to estimate the value of each state in order to provide more useful reinforcement feedback to the actor. The result is two interacting adaptive processes: the actor adapts to the critic, while the critic adapts to the actor. The foundations of actor/critic learning systems date back at least to Samuel's checker program in the late 1950s (Samuel,196...

Read the paper · More papers on PaperTik