Learning long-term dependencies is not as difficult with NARX recurrent neural networks
Tsung-Nan Lin, Bill G. Horne, Peter Tiňo, Clyde Lee Giles · University Libraries (University of Maryland) · 1995
It has recently been shown that gradient descent learning algorithms for recurrent neural networks can perform poorly on tasks that involve long--term dependencies, i.e. those problems for which the desired output depends on inputs presented at times far in the past. In this paper we explore the long--term dependencies problem for a class of architectures called NARX recurrent neural networks, which have powerful representational capabilities. We have previously reported that gradient descent learning is more effective in NARX networks than in recurrent neural network architectures that have "hidden states" on problems including grammatical inference and nonlinear system identification. Typically, the network converges much faster and generalizes better than other networks. The results in this paper are an attempt to explain this phenomenon. We present some experimental results which show that NARX networks can often retain information for two to three times as long as conventional rec...