DESIGNING (APPROXIMATE) OPTIMAL CONTROLLERS via DHP ADAPTIVE CRITICS & NEURAL NETWORKS

George G. Lendaris, Thaddeus T. Shannon · 1999

The objective of this chapter is to provide the reader some guidance in applying the Dual Heuristic Programming (DHP) method in the context of designing neural-network controllers. DHP is a member of the class of Critic methods, which in turn is a member of the class of Reinforcement Learning methods. Development of the DHP method benefited from the confluence of several other developments; the following subsections describe associated background ideas useful in appreciating the DHP method. Subsequent sections will describe the DHP method itself, provide suggestions for application of DHP, and present worked-out examples. 1.1 Learning Algorithms A key distinguishing feature of the computational paradigm known as neural networks is its attribute of attaining knowledge via interaction with its environment (vs. having knowledge programmed in). A significant area of research in the neural network (NN) field has been, and continues to be, that of developing strategies by which the NN accomplishes this extraction of knowledge from its environment-- these are typically referred to as learning algorithms. These algorithms fall into three general categories (in the following descriptions, the terms ‘pupil ’ and ‘teacher ’ designate, respectively, the NN that is learning, and the process used to accomplish the learning; further, except where indicated, the pupil NN is considered an input/output devise): 1. Supervised Learning: This category entails the teacher role having at its disposal full knowledge of the problem context (about which the pupil NN is to learn), and in particular, has available a collection of data pairs (comprising input and associated desired output) with which to conduct the pupil’s learning process.

Read the paper · More papers on PaperTik