Complex Dual-Path Conformer and Convolution Recurrent Network to Speech Enhancement
Xinyu Hao, Zhongdong Wu · 2024
In response to the problem that deep complex convolution recurrent network (DCCRN) cannot effectively model speech features when processing long speech sequences in the time-frequency (T-F) domain, and has low parallelism, a complex dual-path Conformer is proposed to replace the complex Long short term memory (LSTM) module in the baseline model DCCRN, and obtain a new model for T-F domain speech enhancement. The new model inputs noisy speech signals in the time domain, which are then transformed into T-F domain through short-time Fourier transform (STFT). Then the complex spectrogram is gave to the complex encoder to obtain a feature map. Through complex dual-path Conformer, the features of longer speech sequences are modeled to comprehensively utilize global contextual information. Finally, the enhanced speech is obtained through a complex decoder. Experiments were conducted on the public datasets Voice Bank, and the results showed that the proposed model used 2.4M parameters, with a PESQ score of 3.02, an increase of 22.83% compared to DCCRN.