Improved Speech Separation via Dual-Domain Joint Encoder in Time-Domain Networks
Lan Wang, Hai‐Tao Zhang, Youli Qiu, Yanji Jiang, Hao Dong, Pengfei Guo · 2024
Time-domain algorithms have been validated to exhibit excellent performance in speech separation tasks. However, a single time-domain encoding feature is insufficient to comprehensively capture the necessary characteristics for effective speech separation. This study introduces an enhanced speech separation model that refines the time-domain encoder of the classical Conv-TasNet. To address the limitations of conventional time-domain encoders in feature extraction, which inherently impose an upper limit on separation accuracy, this study presents enhancements to the time-domain encoder. A novel dual-domain joint encoding module is devised to incorporate both time-domain and frequency-domain information, thereby bolstering the feature encoding capacity of the separation model. Experimental results demonstrate that, compared to the baseline model, the proposed model achieves improvements of 0.5896 dB and 0.5454 dB in SI-SNRi, and 0.5585 dB and 0.577 dB in SDRi on the WSJ0-2mix and Libri2Mix open-source datasets, respectively. Furthermore, the proposed model surpasses the baseline model in terms of both PESQ and STOI metrics, confirming the efficacy of the dual-domain joint encoder in enhancing separation accuracy and quality.