WnD-UNET: A Waveform and Discrete Wavelet Coefficient-Based 1D Deep Learning Model for Single-Channel Noisy Speech Enhancement
Asaduzzaman, Angkon Deb, Anik Biswas, Rudra Roy, Asif Islam, Celia Shahnaz · 2024
Speech enhancement aims to improve the quality and intelligibility of noisy speech signals. Conventional approaches often struggle with generalizing across diverse noise types and varying signal-to-noise ratios (SNRs), especially in challenging real-world environments. These methods typically rely on hand-crafted features, which may not fully capture the complexities of noisy speech. This research proposes WnD-UNET, a deep learning model combining waveform data with Discrete Wavelet Transform (DWT) coefficients using the db1 wavelet for speech enhancement. The novelty of the proposed approach lies in the concatenation of both time-domain and wavelet-domain features, allowing the model to leverage multi-scale information for more accurate noise suppression. The model significantly outperforms conventional methods, delivering enhanced speech quality and intelligibility across super noisy conditions like -10dB. The proposed algorithm attains a STOI of 0.97 under 5dB noise and 0.79 under -10dB noise.