Estimation of Output SI-SDR of Speech Signals Separated From Noisy Input by Conv-Tasnet

Satoru Emura · 2024

The effectiveness of noisy speech separation in single-channel scenarios depends on various factors such as the signal-to-noise ratio and noise type. In real-world scenarios where clean reference signals cannot be obtained, it is desirable to monitor the effectiveness of the speech separation methods. This study investigates the possibility of monitoring the effectiveness of a neural-network-based speech separation method, Conv-TasNet, using only the separated speech signals. It proposes combining Conv-TasNet with Monte-Carlo dropout and estimating the scale-invariant signal-to-distortion ratio of separated speech signals by employing the inverse of the relative variance of multiple separated signals.

Read the paper · More papers on PaperTik