Time-domain convolutional speech separation network with channel reconstruction and utterance weighting mechanism

Wei Wang, Yi Fan Jiang · 2025

Deep learning based speech separation is currently the main research direction in the time domain. A representative model in this domain is the fully convolutional time-domain audio separation network (Conv-TasNet), which employs dilated convolutions to expand receptive fields. However, this architecture suffers from two primary limitations: First, the fixed convolutional kernels in temporal convolutional networks (TCNs) lack adaptability to input signals, resulting in suboptimal performance when processing signals with varying temporal scales. Second, conventional depthwise separable convolutions exhibit constrained capacity for modeling inter-channel interactions, often neglecting critical correlations between distinct channel features.To address these limitations, we propose a novel architecture that integrates multi-dilated depthwise separable convolution with channel reconstruction and utterance-wise weighting mechanisms. Furthermore, we introduce a coordinate attention (CoordAtt) module into the encoder structure. The CoordAtt mechanism enhances spatial-temporal feature extraction capabilities by enabling the network to adaptively focus on both local and global temporal dependencies. The proposed weighted WD-Conv empowers each convolutional block to dynamically adjust its attention to local features within varying receptive field ranges, while the channel reconstruction unit facilitates efficient inter-channel communication through a synergistic combination of global pooling and channel-wise feature recombination. These innovations collectively enhance the model's suitability for speech separation tasks. Experimental results demonstrate that our enhanced architecture effectively mitigates the inherent deficiencies of Conv-TasNet in speech separation scenarios. Quantitative evaluations confirm significant improvements in separation quality and computational efficiency, ultimately advancing the overall performance of speech separation systems.

Read the paper · More papers on PaperTik