Second-Order Training of Adaptive Critics for Online Process Control

James J. Govindhasamy, Seán McLoone, G.W. Irwin · IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics) · 2005

This paper deals with reinforcement learning for process modeling and control using a model-free, action- dependent adaptive critic (ADAC). A new modified recursive Levenberg Marquardt (RLM) training algorithm, called temporal difference RLM, is developed to improve the ADAC performance. Novel application results for a simulated continuously-stirred-tank-reactor process are included to show the superiority of the new algorithm to conventional temporal-difference stochastic backpropagation.

Read the paper · More papers on PaperTik