Learning with delayed reinforcement in an exploratory probabilistic logic neural network

Catherine E. Myers · Spiral (Imperial College London) · 1990

This thesis is concerned with neural network learning systems which:• learn from the results of their actions, without a distinct training phase (exploratory learning) • require only global evaluations of success (reinforcement learning) • can cope with reinforcements which arrive only after some delay (delay learn ing) First, the thesis develops a node model based on Aleksander's Probabilistic Logic (PLN) node, and which can be used for exploratory reinforcement learning.Like the PLN, it has a training algorithm which requires only a global evaluation of the success of a network output, rather than explicit provision of what the optimal action would have been.Unlike the simple PLN, it allows the training algorithm to be incremental.Therefore behaviours may be shaped gradually and on-line, and separate train and run phases are not required.Next, Attention-Driven Buffering (ADB) is introduced as a means for perform ing delay learning.ADB systems can learn to predict results which occur over inde finitely long delays while maintaining a small number of past states buffered locally at the nodes.The principal idea is that states should be allocated buffer space based not only on recency but also on unpredictability.Unpredictable inputs are assigned a high "Attention", and this increases their ability to compete for buffer space.An ADB system applied to a state traversal task can learn the results of actions in a given state; it can also learn not to enter states with immediate positive reinforce ment, but which lead inevitably to negative reinforcement.Finally, an ADB system has been constructed to perform delay learning tasks resembling those to which an animal such as the octopus can be trained.The sys

Read the paper · More papers on PaperTik