Using Perceptual Quality Features in the Design of the Loss Function for Speech Enhancement

Nicholas Eng, Yusuke Hioka, Catherine Inez Watson · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

Although deep learning has shown success in the domain of speech enhancement, there is still much research that can be undertaken in finding alternative loss functions to train the speech enhancement models. An approach is to utilise features in objective quality metrics, which aim to judge the quality of enhanced or transmitted speech, as part of the loss function. One such objective quality metric utilises perceptual quality dimensions, which identify perceptual quality features that can be derived from the signal and have shown to contribute to perceptual speech quality. In this study, we investigate two such features, cepstrum statistics and MFCC statistics, to be used alongside a baseline loss function for a DNN speech enhancement method. The results from experimentation show that addition of cepstrum statistics to the loss function is detrimental to the scores of objective quality metrics compared to a baseline loss function, however using MFCC statistics in the loss function improves speech quality scores.

Read the paper · More papers on PaperTik