Time Domain Speech Enhancement Using SNR Prediction and Robust Speaker Classification

Jigang Ren, Qirong Mao, Hongjie Jia, Jing‐jing Chen · 2020

Recently, deep neural networks (DNNs) have been successfully used for speech enhancement. Most of them are single-task models, and the extracted features are limited by the speech enhancement model. To extract more speech information, we propose a muti-task learning method for speech enhancement using signal to noise ratio(SNR) prediction. SNR prediction task extracts features of the relationship between noise and speech that cannot be obtained by the speech enhancement model. In addition, we added the task of robust speaker recognition to extract speaker features. By sending SNR features and speaker features to the speech enhancement task, the model can obtain more speech-related information to improve the effect of speech enhancement. Experimental results on a public dataset show that our method achieves the state-of-the-art performance and also gets good scores in terms of subjective quality in time-domain speech enhancement.

Read the paper · More papers on PaperTik