Limitations of gradient methods in sequence learning
Diego Federici · 2002
Recurrent neural networks, such as the well-known Simple recurrent Network (SRN, J. Elman, 1990), trained to predict their own next input vector offer a promising framework for developing internal representations of environmental structure. Current training techniques focus on the use of different gradient methods and genetic search. These techniques have the advantage of being general, to develop distributed representations and to achieve holistic computation. On the other side their generality does not pay off in terms of learning speed, accuracy or flexibility. In this paper a temporal learning problem is analyzed with respect to traditional online learning approaches. The results show that gradient methods do not offer a way to identify and correct the actual cause of misclassifications and so are prone to be stuck on local maxima.