DCCRGAN: Deep Complex Convolution Recurrent Generator Adversarial Network for Speech Enhancement

Huixiang Huang, Renjie Wu, Jingbiao Huang, Jucai Lin, Jun E. Yin · 2022

Generative adversarial network (GAN) based speech enhancement (SE) methods still exist some problems. Some GAN-based systems adopt the same structure from Pixel-to-Pixel directly without special optimization. The importance of the generator network has not been fully explored. Other related researches change the generator network and operate in the time-frequency domain, which ignores the importance of the phase component. In this paper, the real and imaginary components of noisy spectrogram are trained simultaneously. In order to train the complex target effectively, a deep complex convolution recurrent GAN (DCCRGAN) structure is proposed. The complex module builds the correlation between real and imaginary components of noisy spectrogram and has been proved to be effective. For the first time that a complex value network been introduced in GAN-based SE task and the performance of the complex GAN-based network been discussed adequately. Different LSTM layers are used in the generator network to sufficiently explore the speech enhancement performance of DCCRGAN. The experimental results confirm that the proposed DCCRGAN outperforms the state-of-the-art GAN-based SE systems.

Read the paper · More papers on PaperTik