Symbiotic Evolution of Neural Networks in Sequential Decision Tasks
David E. Moriarty · 1997
viii Chapter 1 Introduction 1 1.1 Sequential Decision Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.1.1 Examples of Sequential Decision Tasks . . . . . . . . . . . . . . . . . . . . . . 3 1.1.2 Problem Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.1.3 Properties of Sequential Decision Tasks . . . . . . . . . . . . . . . . . . . . . 5 1.2 Concluding Remarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 Chapter 2 Learning Decision Strategies from Reinforcements 7 2.1 Reinforcement Learning vs. Supervised Learning . . . . . . . . . . . . . . . . . . . . 8 2.2 Temporal Difference Reinforcement Learning . . . . . . . . . . . . . . . . . . . . . . 8 2.2.1 Learning Through Temporal Differences . . . . . . . . . . . . . . . . . . . . . 9 2.2.2 The Adaptive Heuristic Critic . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.2.3 Q-learning . . . . . . . . . . . . . . . . . . . . . . . . . . . ....