Reinforcement learning neurocontroller applied to a 2-DOF manipulator
Marco Pérez‐Cisneros, R.R. Leal Ascencio, Peter A. Cook · 2002
This paper describes the capabilities of a reinforcement learning (RL) algorithm which uses two neural net structures to produce a direct inverse neurocontrol scheme. The Pendubot/sup TM/ a double inverted pendulum which is a nonlinear dynamic system inherently unstable, is used as benchmark plant because it could be attractive for testing control schemes. The RL neurocontroller is a learning system which consists of two connectionist nets, the action net (AN) and the evaluation net (EN). The action net generates the system's behavior and the evaluation net learns an evaluation function of Pendubot's states. A zero magnitude force is not permitted and the neurocontroller always is supplying a control signal for Pendubot/sup TM/. The paper also describes a neurocontroller feature consisting of a nonlinear function added to the control signal when the second link approaches critic states. Results from training and operation stage are summarized for both neurocontrollers. Finally, the mass of link 2 of Pendubot/sup TM/ is altered increasing and decreasing its magnitude in order to observe the generalization capabilities in the neurocontroller. This last experiment is also documented.