Efficient Riemannian training of recurrent neural networks for learning symbolic data sequences

Yann Ollivier · arXiv (Cornell University) · 2013

Recurrent neural networks are powerful models for sequential data, able to represent complex dependencies in the sequence that simpler models such as hidden Markov models cannot handle. Yet they are notoriously hard to train. Here we introduce an effective training procedure using a gradient ascent in a metric inspired by Riemannian geometry: this produces an algorithm independent from design choices such as the encoding of parameters and unit activities. This metric gradient ascent is designed to have an algorithmic cost close to backpropagation through time for sparsely connected networks. We also introduce \emph{persistent contextual neural networks} (PCNNs) as a variant of recurrent neural networks. PCNNs feature an architecture inspired by finite automata and a modified time evolution to better model long-distance effects. PCNNs are demonstrated to effectively capture a variety of complex algorithmic constraints on hard synthetic problems: basic block nesting as in context-free grammars (an important feature of natural languages, but difficult to learn), intersections of multiple independent Markov-type relations, or long-distance relationships such as the distant-XOR problem. On this problem, PCNNs perform better than more complex state-of-the-art algorithms. Thanks to the metric update, fewer gradient steps and training samples are needed: for instance, a generating model for sequences of the form $a^nb^n$ can be learned from only 10 samples in under two minutes, even with $n$ ranging in the thousands.

Read the paper · More papers on PaperTik