Speeding up deep neural network based speech recognition systems

Yeming Xiao · Journal of Software · 2014

Abstract—Recently, deep neural network (DNN) based a-coustic modeling has been successfully applied to large vocabulary continuous speech recognition (LVCSR) tasks. A relative word error reduction around 20 % can be achieved compared to a state-of-the-art discriminatively trained Gaussian Mixture Model (GMM). However, due to the huge number of parameters in the DNN, real-time de-coding is a bottleneck for the DNN based speech recognition systems. In this paper, we adopt several techniques for the speed optimization of the DNN-based system. Specifically, we use singular value decomposition (SVD) to reduce the model parameters, use the SSE instruction sets for the parallel calculation in the data space, and quantize the model parameters reasonably to convert the floating-point arithmetic into fixed-point arithmetic. Besides, taking the characteristics of speech signal into account, we use a frame-skipping method when evaluating the posterior probabilities. Finally, compared to the un-optimized baseline system, with negligible recognition performance loss, the decoding real-time factor of the optimized one is significantly reduced, from 6.1 to 0.31. And this response speed can basically meet the requirement of our real applications. Index Terms—Large Vocabulary Continuous Speech Recog-

Read the paper · More papers on PaperTik