Multi-Stream Diffusion Model for Probabilistic Integration of Model-Based and Data-Driven Speech Enhancement
Tomohiro Nakatani, Naoyuki Kamo, Marc Delcroix, Shoko Araki · 2024
This paper introduces a novel approach that leverages a diffusion model to seamlessly integrate multiple Speech Enhancement (SE) methods, resulting in high-quality speech estimates. Traditionally, researchers have combined various strategies, including model-based signal processing and data-driven neural networks (NNs), to achieve state-of-the-art SE performance. However, these methods often relied on manually designed integration schemes, which lacked optimality. In contrast, we propose a diffusion model that optimizes the integration of SE methods in a probabilistic manner using training data. The proposed method is called a multi-stream Score-based Generative Model for SE (mSGMSE). We experimentally demonstrate that our approach significantly enhances speech quality by integrating a diffusion model-based SE with a Complex Spectral Mapping (CSM)-based SE and Weighted Prediction Error (WPE) dereverberation. Utilizing a two-microphone setup, our method reduces the Word Error Rate (WER) obtained by the ESPnet2 automatic speech recognizer from 6.14% to 3.46% on the REVERB challenge real dataset.