Elimination of Noise in Speech Signals Utilizing METS Deep Convolutional Neural Networks
Jung-Hua Wang · JOURNAL OF HIGH-FREQUENCY COMMUNICATION TECHNOLOGIES · 2025
In the realm of digital signal processing, speech enhancement plays a crucial role in applications such as teleconferencing, voice recognition, and biometric systems. Noise and distortions significantly affect speech quality, necessitating advanced enhancement techniques. This paper proposes an optimized deep convolutional neural network (CNN)-based speech enhancement method, integrating signal subspace searching and the minimum error and time-spectral estimator (METS). The model is trained and evaluated using the LJ Speech Dataset, augmented with various noise conditions. Experimental results demonstrate that the proposed method achieves a PESQ of 3.7, STOI of 0.92, and SNR improvement of 12.3 dB, outperforming traditional and deep learning-based methods such as Spectral Subtraction, Wiener Filtering, MMSE, SEGAN, and DCRN. The integration of METS refines the spectral estimation, while CNN effectively reconstructs speech features, leading to better intelligibility and reduced spectral distortion. Future research will focus on real-time processing and adaptive noise handling, ensuring robust speech enhancement for diverse applications.