Cptnn: Cross-Parallel Transformer Neural Network For Time-Domain Speech Enhancement

Kai Wang, Bengbeng He, Wei‐Ping Zhu · 2022

In this paper, we propose a novel cross-parallel transformer neural network (CPTNN) for end-to-end speech enhancement in the time domain. The new structure is comprised of an encoder, a cross-parallel transformer module (CPTM), a masking module and a decoder. The encoder first maps the input waveform of noisy speech into feature representations. The CPTM consists of four residually connected cross-parallel transformer blocks, each utilizing local and global transformers to simultaneously extract local and global features which are then fused by a cross-attention based transformer to obtain a better contextual feature representation. The masking module generates a mask to multiply with encoder output, producing the masked encoder features which will be finally used for reconstructing the enhanced speech by the decoder. Experiments are undertaken on the benchmark dataset, indicating that our CPTNN achieves a better performance than state-of-the-art methods in terms of most evaluation criteria while maintaining the lowest model parameters.

Read the paper · More papers on PaperTik