Modifying Flow Matching for Generative Speech Enhancement

Roman Korostik, Rauf Nasretdinov, Ante Jukić · 2025

Diffusion-based generative models have been shown to be highly effective in various speech enhancement tasks. This work presents an analysis of a flow matching-based framework for generative speech enhancement as a simpler alternative to diffusion. Four different modifications to flow matching are proposed, employing an informed prior, a data prediction loss, deterministic inference, and early stopping. The proposed variants are evaluated on speech denoising, demonstrating performance comparable to a previous state-of-the art model using the same data setup. Through ablation studies, an efficient deterministic one-step inference configuration is proposed, which does not require any advanced training techniques such as pre-training or distillation. The proposed variants are also evaluated on speech dereverberation, demonstrating that stochastic inference without informed prior is preferable for this task.

Read the paper · More papers on PaperTik