A Comparative Study of Time and Frequency Domain Approaches in Single Channel Foreground Speech Enhancement using Deep Neural Networks
Debabrata Gogoi, Sushanta Kabir Dutta · 2024
Speech enhancement plays a crucial role in a large number of applications, by aiming to improve the intelligibility and quality of speech signals degraded due to the presence of noise. This paper investigates the performance of deep learning inspired architectures with Convolutional Neural Networks (CNNs) for speech enhancement using two different input representations: time series and log-Mel spectrograms of data. Using a suboptimal dataset with limited size, limited training epochs, and constrained hardware, we evaluate the efficacy of these representations under the influence of some random noises. The performances of both models developed accordingly were evaluated on same the test set. The model using time series data as input demonstrated better performance in capturing local patterns and effectively reducing stationary noise, while the one using log-mel spectrogram as input, excelled in modeling the temporal dynamics of the speech signal and therefore reduced the non-stationary noise in the data.