Deep Learning Based Automatic Noisy Speech Classification for Enhanced Speech Analysis
K Sankavi, G. Jyothish Lal, B. Premjith · 2024
In real-life scenarios, the speech signals are often corrupted by various background noises such as vehicles, trains, fans, wind, rain, air-conditioners, and machinery sounds. Hence, ensuring the quality of speech signals is most important for the efficient performance of speech and audio processing systems. In this paper, we study the performance of deep learning-based noisy speech classification (NSC) methods for noise-specific speech denoising or speech enhancement for better noise reduction and preserving speech intelligibility and naturalness. We explore the performance of the autocorrelation function and spectrogram features in combination with 1D and 2D convolutional neural networks (CNNs) and long short-term memory (LSTM) methods. The proposed NSC method is evaluated using a large dataset comprising noise free and noisy speech signals corrupted with five types of noise at different noise amplitude levels. Experimental results demonstrate that the spectrogram-based LSTM method achieves an overall accuracy (ACC) 99.60% with a model size of 11.6 MB and a computational time of 0.47 ms. The huge computational complexity reduction using the lightweight NSC methods can improve the energy efficiency and battery life of edge speech analysis systems.