SE-DPTUNet: Dual-Path Transformer based U-Net for Speech Enhancement

Bengbeng He, Kai Wang, Wei‐Ping Zhu · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

In this paper, we propose a novel transformer-based U-Net model for end-to-end speech enhancement in the time domain, called SE-DPTUNet, which is comprised of an encoding module, a proposed dual-path transformer module and a decoding module. The encoding module extracts low-level features from input noisy speech. The dual-path transformer module exploits transformer encoder-decoder layers to model long-range speech sequence by extracting short- and long-term information of speech sequences, generating the contextual features. Meanwhile, the down- and up-sampling blocks are adopted in each encoder-decoder layer to obtain the hierarchical features. The decoding module is applied to reconstruct the enhanced speech based on the output from transformer module. Experimental results on the benchmark dataset indicate that our proposed SE-DPTUNet achieves a competitive performance compared with existing models while having significantly reduced model complexity.

Read the paper · More papers on PaperTik