Speech Enhancement Method Based On LSTM Neural Network for Speech Recognition
Ming Liu, Yujun Wang, Jin Wang, Jing Wang, Xiang Xie · 2018
Long Short-Term Memory (LSTM), a special kind of Recurrent Neural Network (RNN), is capable of learning long-term dependencies. In this paper, a kind of speech enhancement method is proposed for LSTM network structure to cope with the speech features, with the purpose of improving the speech recognition rate. This method utilizes the LSTM structure in reference to the acoustic model and crossover residual network to construct the front-end enhancement module. We trained and compared DNN, CNN, LSTM and BLSTM models with various numbers of parameters. The experimental results show that, the LSTM model performs the best in the test set and the real scene. The noise reduction effects are the best when the noise is reduced from 31.23% to 25.89% on the Xiaomi speaker test set1.