Joint optimization of modified ideal radio mask and deep neural networks for monaural speech enhancement

Wei Han, Congming Wu, Xiongwei Zhang, Qiye Zhang, Songting Bai · 2017

Monaural speech enhancement is a key yet challenging problem in speech area, which is always used as a pre-processing step of robust speech processing. Deep learning has proved to be very successful for solving this issue. In this paper, a new approach for enhancing the noisy speech in a single channel recording is presented. We propose a modified ideal ratio mask (IRM) which calculated by normalized cross-correlation coefficient (NCCC). Then, we jointly optimize the modified ideal ratio mask and deep neural network (DNN). Evaluation experiments on TIMIT database with 20 noise types at different signal-to-noise ratio (SNR) situations demonstrate the effectiveness of the proposed approach compared with the reference DNN-based enhancement approaches, no matter whether the noise matched the training set or not.

Read the paper · More papers on PaperTik