DPTCN-ATPP: Multi-scale End-to-end Modeling for Single-channel Speech Separation
Yanmin Zhu, Xiang Zheng, Xinrong Wu, Wanning Liu, Lei Pi, Meiju Chen · 2021
Dual-Path RNN(DPRNN) has achieved great progress in single-channel speech separation. However, RNN-based model needs to pass information through intermediate states and it does not allow parallel computing. Meanwhile, inter-chunk modeling of DPRNN only modeled between multiples of the chunk length, which means the underutilization of contextual information. To address these problems, we propose a multi-scale dual-path temporal convolutional network, DPTCN-ATPP. DPTCN-ATPP utilizes a stacked of one-dimensional dilated convolutions (1-D CNNs) instead of RNN to learn local and global information. Besides, DPTCN-ATPP applies parallel encoders to extract multi-scale input features. Finally, DPTCN-ATPP introduces the ATPP block to summarize multi-scale deep features. The experiment results show that DPTCN-ATPP achieves 19.6dB and 19.8dB on the metrics of scale-invariant source-to-noise ratio(SI-SNR) and source-to-distortion (SDR) respectively, outperforming baseline DPRNN and single-scale TCN (Conv-TasNet).