Learning to predict a context-free language: analysis of dynamics in recurrent hidden units
Mikael Bodén · 1999
Recurrent neural network processing of regular language is reasonably well understood. Recent work has examined the less familiar question of context-free languages. Previous results regarding the language a suggest that while it is possible for a small recurrent network to process context-free languages, learning them is di#cult. This paper considers the reasons underlying this di#culty by considering the relationship between the dynamics of the network and weightspace. We are able to show that the dynamics required for the solution lie in a region of weightspace close to a bifurcation point where small changes in weights may result in radically di#erent network behaviour. Furthermore, we show that the error gradient information in this region is highly irregular. We conclude that any gradient-based learning method will experience di#culty in learning the language due to the nature of the space, and that a more promising approach to improving learning performance may be to make weight changes in a nonindependent manner.