Robust Speech Enhancement Using Convolution-Recurrent Frameworks and Wavelet Pooling
G V Soumya, R Paramesha, M. Poornima, Kiran Puttegowda, B Susmitha, Sunil Kumar D S, B D Parameshachari · 2024
Speech enhancement is a crucial and challenging task in many applications. In this work, we present an end-to-end data-driven system for enhancing the quality of speech signals using a convolutional-recurrent neural network. We present a quantitative and qualitative analysis of our speech enhancement system on a real-world noisy speech dataset and evaluate our proposed system's performance using several metrics such as Signal-to-Noise Ratio (SNR), Perceptual Evaluation of Speech Quality (PESQ), Short-time Objective Intelligibility (STOI). We have employed wavelet pooling mechanism instead of max-pooling layer in the convolutional layer of our proposed model and compared the performances of these variants. Based on our experiments, we demonstrate that our model's performance on noisy speech signals using Haar wavelet is better than when using max pooling. In addition, wavelet-based approach results in faster convergence during training as compared to other variants. Moreover, we observed that the wavelet-based approach facilitates faster convergence during training, highlighting its efficiency in optimizing the learning process. The results of the study demonstrate that the proposed speech enhancement system, which utilizes a convolutional-recurrent neural network (CNN-RNN) and wavelet pooling, significantly improves the performance of noisy speech signal enhancement. These findings underscore the potential of wavelet pooling as a powerful alternative to conventional pooling methods, particularly in tasks involving complex temporal and spectral structures, such as speech enhancement.