Noise-Robust Speech Enhancement via Deep Learning Models
Pasin Israsena · 2025
This paper presents a novel deep learning approach for speech enhancement that accounts for varying environmental noise conditions. The proposed method utilises a convolutional neural network (CNN) architecture with separately trained models tailored to different noise environments. The CNN operates on a spectrum-based representation and is trained to estimate Ideal Ratio Masking (IRM). The effectiveness of the proposed technique was evaluated using objective metrics, including the Perceptual Evaluation of Speech Quality (PESQ) score and the Short-Time Objective Intelligibility (STOI) score. When compared to noisy baseline conditions and alternative deep learning methods that do not account for environmental classifications, the proposed approach demonstrated statistically significant improvements, with performance gains of up to 10.39%.