SEformer: Dual-Path Conformer Neural Network is a Good Speech Denoiser

Kai Wang, Dimitrios Hatzinakos · 2023

In this paper, we propose the SEformer, an efficient dual-path conformer neural network for speech enhancement. The proposed SEformer is comprised of an encoder, a decoder and dual-path conformer blocks in between. First, the encoder receives the noisy speech waveform to generate the speech features. Then, each proposed dual-path conformer block employs a temporal and frequency conformer in parallel to simultaneously extract temporal and frequency features of speech sequences, which are then fused by a transformer block for contextual representation. Finally, the decoder transforms the contextual speech features into a mask to filter out noise-related information of noisy speech, generating the estimated clean speech. Experimental results on the benchmark dataset exhibit that our SEformer yields a competitive performance to existing state-of-the-art methods while containing the fewest model parameters (about 590k).

Read the paper · More papers on PaperTik