A Two-Stage Speech Enhancement Method Based on CNNLMS
Xin Zheng, Ying He, Chong Zhang · 2023
This letter presents a two-stage speech enhancement framework based on a one-dimensional convolutional neural network (1D-CNN) and the Least Mean Square (LMS) algorithm. Feature extraction and channel filtering can reduce background noise and enhance speech quality. Combining these techniques can improve speech quality in low signal-to-noise ratio (SNR) conditions. Temporal features are extracted with a one-dimensional convolutional neural network. The nonlinear mapping between the signals is estimated. The output signals of the neural network are further denoised using the LMS algorithm, which can adjust the filter coefficients to suppress noise signals and improve the clarity of speech signals. Compared with other algorithms, such as a neural network algorithm based on spectral mapping, or the algorithm based on time-frequency masking, our proposed CNNLMS method can increase the Short-Time objective intelligibility (STOI) value of speech by 0.076. The perceptual evaluation of speech quality (PESQ) score is increased by 1.2. Even in situations where the actual environment does not match the training environment, this algorithm can still maintain effective speech enhancement performance.