Enhancing Single-Channel Speech Processing With Advanced Noise-Reduction Techniques
D Deepa, C. S., V P Shreyaa, Thangavel · 2025
This paper investigates the performance of a CNN-based U-Net architecture to enhance the quality of speech signals degraded by noise. The goal here is to measure how much clarity and intelligibility of speech improves for children who might be having difficulty perceiving speech because of background urban noise. An extensive experiment was performed on a dataset which had clean speech from RAVDESS and the urban background from UrbanSound8K.Using the Short-Time Fourier Transform (STFT), the audio samples were changed to time frequency representation and were applied to a U-Net network. Various objective measures were utilized to evaluate improvements such as the signal to noise ratio (SNR), Itakura-Saito distance, RMSE, or the short-time objective intelligibility (STOI). They were well contrasted with standard filtering techniques such as Wiener filtering or Weinberg masking. The improvements in these technologies were witnessed with an increase in the STOI score which increased from 0.71 to 0.83 which indicates an improvement in the speech quality. In addition, using the Mean Opinion Score (MOS) which asks listeners to provide their qualitative assessments of the enhanced audio signals in terms of certain topics. With the discussions, it was confirmed that U-Net is useful in speech improvement and quality enhancement through noise reduction. The experiments revealed that the performance of the fuse deep CNN U-Net architecture is superior to legacy approaches for restoring the quality of speech in the presence of background noise. These insights may inform the design and development of changing amplifiers and other audio systems to assist communication in ‘difficult’ listening conditions.