Integrating Advanced Convolutional and Recurrent Architectures for Speech Enhancement
G. Jenifa, R M Sai Likhita, M Vineela, T V Vishnu Vardhan Babu, Shaik Sabiya · 2025
This paper addresses the challenge of enhancing speech clarity and intelligibility in noisy environments using time domain deep learning models. Traditional approaches like spectral subtraction and Wiener filtering struggle with non-stationary noise and often introduce artifacts. To overcome these limitations, we evaluate the performance of UNet, Recurrent Neural Network (RNN), and Long Short Term Memory (LSTM) architectures for direct waveform based speech enhancement. Unlike frequency domain models, our approach processes raw signals, preserving temporal dynamics and minimizing distortion. Experimental results demonstrate that the UNet model outperforms RNN and LSTM, achieving a 97.47% improvement in SNR, an 18.48% gain in PESQ, and a 2.10% increase in STOI. These findings underscore the effectiveness of time domain deep learning models for robust and intelligible speech enhancement.