Estimation of Output SI-SDR Solely from Enhanced Speech Signals in Diffusion-Based Generative Speech Enhancement Method

Satoru Emura · 2024

Recently, a new class of generative models, diffusion-based generative models (DGMs), has been introduced for speech enhancement. The effectiveness of speech enhancement depends on various factors, such as the signal-to-noise ratio and noise types. In real-world scenarios where clean reference signals cannot be obtained, it is desirable to monitor the effectiveness of speech enhancement methods. This study investigates the possibility of monitoring the effectiveness of the DGM-based speech enhancement, using only the enhanced speech signals. It proposes estimating the scale-invariant signal-to-distortion ratio of enhanced speech signals by employing the inverse of the relative variance of multiple enhanced signals.

Read the paper · More papers on PaperTik