Diffusion Model-Based MIMO Speech Denoising and Dereverberation
Rino Kimura, Tomohiro Nakatani, Naoyuki Kamo, Marc Delcroix, Shoko Araki, Tetsuya Ueda, Shoji Makino · 2024
This paper presents an extension of a diffusion model-based single microphone Speech Enhancement (SE) method, known as the Score-Based Generative Model for SE (SGMSE), to a Multi-Input Multi-Output (MIMO) SE. The extended method is called a multi-stream SGMSE (mSGMSE). MIMO SE’s goal is to estimate multi-microphone clean speech signals with spatial cues from noisy, reverberant speech signals captured by a distant microphone array. mSGMSE models the conditional distribution of clean speech signals given the captured signals for multi-microphone signals using a diffusion model and generates clean speech estimates by using the reverse diffusion process. We also propose techniques to make mSGMSE computationally efficient and adaptable to various or unknown array geometries for general MIMO SE scenarios. Experiments show that mSGMSE outperforms SGMSE (when separately applied to each microphone signal) in terms of signal quality, spatial cues, and computation times. mSGMSE also significantly improves the automatic speech recognition performance when applied to the REVERB challenge real dataset, which has substantial mismatches including array geometries from the training dataset. Finally, we confirm that Weighted Prediction Error dereverberation (WPE) preprocessing can further enhance mSGMSE more than SGMSE.