Are Recent Deep Learning-Based Speech Enhancement Methods Ready to Confront Real-World Noisy Environments?
Candy Olivia Mawalim, Shogo Okada, Masashi Unoki · 2024
Recent advancements in speech enhancement techniques have ignited interest in improving speech quality and intelligibility.However, the effectiveness of recently proposed methods is unclear.In this paper, a comprehensive analysis of modern deep learning-based speech enhancement approaches is presented.Through evaluations using the Deep Suppression Noise and Clarity Enhancement Challenge datasets, we assess the performances of three methods: Denoiser, DeepFilterNet3, and FullSubNet+.Our findings reveal nuanced performance differences among these methods, with varying efficacy across datasets.While objective metrics offer valuable insights, they struggle to represent complex scenarios with multiple noise sources.Leveraging ASR-based methods for these scenarios shows promise but may induce critical hallucination effects.Our study emphasizes the need for ongoing research to refine techniques for diverse real-world environments.