On the Efficiency of Recurrent Neural Network Optimization Algorithms
Ben Krause, Liang Lu, Iain Murray, Steve J. Renals · Edinburgh Research Explorer (University of Edinburgh) · 2015
This study compares the sequential and parallel efficiency of training Recurrent Neural Networks (RNNs) with Hessian-free optimization versus a gradient descent variant. Experiments are performed using the long short term memory (LSTM) architecture and the newly proposed multiplicative LSTM (mLSTM) architecture. Results demonstrate a number of insights into these architectures and optimization algorithms, including that Hessian-free optimization has the potential for large efficiency gains in a highly parallel setup.