Long short-term memory in recurrent neural networks
Felix A. Gers · 2001
For a long time, recurrent neural networks (RNNs) were thought to be theoretically fascinating. Unlike standard feed-forward networks RNNs can deal with arbitrary input sequences instead of static input data only. This combined with the ability to memorize relevant events over time makes recurrent networks in principal more powerful than standard feed-forward networks. The set of potential applications is enormous: any task that requires to learn how to use memory is a potential task for recurrent networks. Potential application areas include time series prediction, motor control in non-Markovian environments and rhythm detection (in music and speech). Previous successes in real world applications, with recurrent networks were limited, however, due to practical problems when long time lags between relevant events make learning dicult. For these applications conventional gradient-based recurrent network algorithms for learning to store information over extended time intervals take too long. The main reason for this failure is the rapid decay of back-propagated error. The \\Long Short Term Memory" (LSTM) algorithm overcomes this and related problems by enforcing constant error ow. Using gradient descent, LSTM explicitly learns when to store information and when to access it. In this thesis we extend, analyze, and apply the LSTM algorithm. In particular, we identify two weaknesses of LSTM, oer solutions and modify the algorithm accordingly: (1) We recognize a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset. Without resets, the state may grow indenitely and eventually cause the network to break down. Our remedy is a novel, adaptiv...