Efficient GPU-based training of recurrent neural network language models using spliced sentence bunch

X. Chen, Y. Wang, Xiaobing Liu, Mark Gales, Philip C. Woodland · 2014

Recurrent neural network language models (RNNLMs) are be-coming increasingly popular for a range of applications includ-ing speech recognition. However, an important issue that limits the quantity of data, and hence their possible application ar-eas, is the computational cost in training. A standard approach to handle this problem is to use class-based outputs, allowing systems to be trained on CPUs. This paper describes an alter-native approach that allows RNNLMs to be efficiently trained on GPUs. This enables larger quantities of data to be used, and networks with an unclustered, full output layer to be trained. To improve efficiency on GPUs, multiple sentences are “spliced” together for each mini-batch or “bunch ” in training. On a large vocabulary conversational telephone speech recognition task, the training time was reduced by a factor of 27 over the stan-dard CPU-based RNNLM toolkit. The use of an unclustered, full output layer also improves perplexity and recognition per-formance over class-based RNNLMs. Index Terms: language models, recurrent neural network, speech recognition, GPU

Read the paper · More papers on PaperTik