Q-Learning with Hidden-Unit Restarting

Charles W. Anderson · 1992

Platt's resource-allocation network (RAN) (Platt, 1991a, 1991b) is modified for a reinforcement-learning paradigm and to "restart" existing hidden units rather than adding new units. After restarting, units continue to learn via back-propagation. The resulting restart algorithm is tested in a Q-learning network that learns to solve an inverted pendulum problem. Solutions are found faster on average with the restart algorithm than without it. 1 Introduction The goal of supervised learning is the discovery of a compact representation that generalizes well. Such representations are typically found by incremental, gradientbased search, such as error back-propagation. However, in the early stages of learning a control task, we are more concerned with fast learning than a compact representation. This implies a local representation with the extreme being the memorization of each experience. An initially local representation is also advantageous when the learning component is operating in par...

Read the paper · More papers on PaperTik