Real-Time Audio Noise Reduction and Speech Enhancement Using LadderNet With Hybrid Spectrogram Time-Domain Audio Separation Network
Ayesha Siddiqua, CH Hussaian Basha, Haider Mohammed Abbas, P. Merlin Sheela, V. Indhumathi · 2024
In recent years, the rapid growth of modern technology has improved people lives, but it also comes with a downside: increased exposure to noise from complex industrial infrastructure and machinery. To overcome this issue, the existing method namely a single-channel Deep Neural Network (DNN) model have developed for speech signal processing, but it led to limitations in online processing. Therefore, this research proposes a LadderNet with Hybrid Spectrogram Time-domain Audio Separation Network (HSTasNet) model for reducing the noise and enhancing speech enhancement in real-time audio processing. This procedure begins with the collection of data from real-time audio signal and then preprocessed with spectral gating filter to reduce noise from signals. After that, t-distributed Stochastic Neighbor Embedding (t-SNE) technique to extract the optimal features from preprocessed signals. Finally, LadderNet and HS-TasNet techniques are employed for noise reduction and speech enhancement. As per the results, the proposed LadderNet-HSTasNet model offered outstanding results in terms of Perceptual Evaluation of Speech Quality (PESQ) of 3.847, Short-Time Objective Intelligibility (STOI) of 1.874, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) of 16.792 when compared with existing Dual-Path Recurrent Neural Network (DPRNN) model.