SPP-DeepFilterNet: Integrating Speech Presence Probability for Speech Enhancement using Deep Learning
Nitya Tiwari, K. S. Nataraj, Vimal Prasad · 2024
Speech enhancement, critical for applications like voice communication, hearing aids, and speech recognition, has greatly advanced with the integration of deep learning and filtering approaches. Hybrid methods combining these techniques have shown exceptional promise in mitigating complex noise sources and significantly improving speech quality. However, the current state-of-the-art deep filter model for speech enhancement faces challenges in achieving optimal performance, particularly in complex acoustic environments with diverse noise sources and variable signal conditions. To address these limitations, this research investigates methods for enhancing the deep filter model’s capabilities by integrating additional signal features. Specifically, the research objective is to enhance the performance of the existing DeepFilterNet2 model through the integration of the speech presence probability (SPP) feature as an additional feature, alongside the existing ERB (equivalent rectangular bandwidth) and complex features. Termed as SPP-DeepFilterNet, our proposed model leverages a speech model comprising both periodic and stochastic components. The model’s first stage operates in the ERB domain, enhancing the speech envelope, while the second stage utilizes deep filtering to enhance the periodic component. We hypothesize that the inclusion of SPP synergistically enhances the model’s ability to improve speech quality and intelligibility, ensuring adaptability and robustness across diverse real-world scenarios. Evaluation using objective measures and spectrogram analysis demonstrates improved performance, particularly in unseen noisy conditions, validating the efficacy of the proposed approach.