Improving the performance of DEMUCS in speech enhancement with the perceptual metric loss

Zong-Tai Wu, Yan-Tong Chen, Jeih-weih Hung · 2022 IEEE International Conference on Consumer Electronics - Taiwan · 2022

This study attempts to revise the loss function used in the well-known method, DEMUCS, to improve its speech enhancement performance. DEMUCS is a celebrated deep learning architecture initially designed for audio source separation, while it can also be applied to speech enhancement. In DEMUCS, the objective function is to minimize the loss function that contains the time-domain waveform loss and the multi-resolution short-time Fourier transform (STFT) loss. Here, we propose to directly adopt the perceptual metric-aware losses for speech enhancement, which involve perceptual speech quality (PESQ) and short-time objective intelligibility (STOI) to learn the DEMUCS network. The preliminary experimental results indicate that when jointly optimizing either of the two perceptual metrics (PESQ, STOI) and minimizing the original time- and spectral-domain losses, the speech enhancement performance of DEMUCS can be further promoted.

Read the paper · More papers on PaperTik