A Voice-Activity-Aware Loss Function for Continuous Speech Separation

Conggui Liu, Yoshinao Sato · 2023

Continuous speech separation (CSS) is gaining popularity as a realistic scenario in which speech utterances partially overlap. Most previous studies on CSS employed the conventional scale-invariant signal-to-noise ratio (SI-SNR) loss, which is ill-defined when the reference audio is silent and thus requires some modifications. Hence, we propose a voice-activity-aware loss function that combines SI-SNR loss for speech segments and logarithmic root mean square loss for nonspeech segments. We trained and evaluated the block-online temporal convolutional network models on synthetic single-channel speech mixtures in noisy and reverberant environments. The results demonstrated that the proposed loss function outperformed the conventional loss function. Furthermore, qualitative analyses indicated that the proposed loss function accurately detected speech onsets and offsets, yielding reduced residual noise in nonspeech segments.

Read the paper · More papers on PaperTik