Non-intrusive Intelligibility Prediction of Noise-suppressed Speech Based on Neural Network

Fuqiang Ye, Zexin Liu, Fei Chen · 2022

While many non-intrusive intelligibility indices have been developed recently, they did not work well for speech containing non-linear distortion, which was largely introduced by noise-suppression algorithms. This motivated the present work to propose a model based on neural network (using its non-linear mapping capacity) to non-intrusively predict the intelligibility of noise-suppressed speech. In this study, Mel-frequency cepstrum coefficients (MFCCs) and dynamic MFCCs were combined as the input features of deep belief network (DBN) to predict the intelligibility of noise-suppressed speech. The predicted scores were correlated with the intelligibility scores obtained from normal-hearing listeners presented with noise-suppressed speech. High correlation ($\mathrm{r}=0.81$), compared to$r=0.37$obtained from the existing benchmark non-intrusive intelligibility metric and$r=0.82$from two widely-used intrusive intelligibility indices, suggested that the DBN-based model may work as an efficient intelligibility predictor of non-linearly distorted speech without access to the clean reference signal.

Read the paper · More papers on PaperTik