Leveraging the perceptual metric loss to improve the DEMUCS system in speech enhancement
Qi-Wei Hong, Chi-En Dai, Hui-Chun Hsu, Zong-Tai Wu, Jeih-weih Hung · 2022
This study aims to improve the source separation technique, DEMUCS, by revising the respective loss function. DEMUCS, developed by Facebook Team, is built on the Wave-U-Net and consists of convolutional layer encoding and decoding blocks with an LSTM layer in between. The applied loss function in DEMUCS contains wave-domain L1 distance and multi-scale short-time-Fourier-transform (STFT) loss.We present to revise the original loss by considering the perceptual metric scores, including perceptual speech quality (PESQ) and short-time objective intelligibility (STOI). The new loss function becomes a weighted sum of the original loss and the losses of STOI and PESQ, hoping to highlight the perceptual quality of the enhanced utterances.According to the preliminary experiments conducted on the VoiceBank-DEMUCS task, the DEMUCS network with the modified loss function provides the noise-corrupted utterances with superior objective perceptual metric scores (PESQ and STOI). These results indicate that the presented work benefits DEMUCS in speech enhancement performance.