Masking based Spectral Feature Enhancement for Robust Automatic Speech Recognition
Chunlei Liu, Longbiao Wang, Jianwu Dang · 2020
The accuracy of the automatic speech recognition (ASR) system will decrease under reverberant conditions. Feature enhancement as the front-end of an ASR system can improve the performance in adverse environments. In this research work, we propose modified masking based deep learning method to enhance the log Mel filterbank coefficients of reverberant speech. We eliminate the zero-value amplitude output by adding an adjustment layer to the DNN, thus improving the accuracy of spectral feature estimation. As the masking method relies on the human hearing mechanism, its feature enhancement effect is superior than the state-of-the-art spectral feature mapping method. ASR experiments with the REVERB 2014 stereo (reverberant and clean) data proved that the proposed method achieved excellent results.