DCE-CDPPTnet: Dense Connected Encoder Cross Dual-Path Parrel Transformer Network for Multi-Channel Speech Separation
Chenjie Zhuang, Lin Zhou, Y. F. Cao, Qirui Wang, Y.M. Cheng · 2024
In recent years, there has been increasing attention given to the end-to-end time domain model for speech separation task, which has demonstrated superior performance. Specifically, research in time domain speech separation has focused on two main areas: extracting more effective features and improving the modeling of temporal speech sequences. To tackle these challenges, we propose the dense connected encoder cross dual-path parallel transformer network (DCE-CDPPTnet) for multi-channel speech separation. By utilizing a dense connected encoder, our network is able to enhance feature extraction through multilayer convolutional layers organized by dense connected structure. Furthermore, this encoder also mitigates the issue of gradient disappearance. In order to better model long-time speech sequences, our proposed model incorporates the cross dual-path parallel transformer. This transformer utilizes both the intra improved transformer and the inter improved transformer to capture local and global information, respectively. Moreover, the CDPPTnet enables local information and global information can interact by parallelizing the intra improved transformer and the inter improved transformer. Simulation results under various model configurations, demonstrate the superior performance of the proposed DCE-CDPPTnet compared to the filter-and-sum network with transform-average-concatenate module (FaSNet-TAC).