A Review of Speech Separation Focusing on TasNet, Conv-TasNet, and DPRNN
Mingyi Liu, Yuzhuo Zhang · 2025
Speech separation technology is crucial for high-quality online audio calls, and well-known challenges like the cocktail party problem have been widely discussed. Recently, deep neural networks (DNNs) have become powerful tools for addressing problems related to voice isolation. Deep learning-based single-channel speech separation models in the time domain have achieved significant success in noise-free scenarios. Based on this, this paper first reviews recent models of speech separation then delves into the architecture and performance of TasNet, a time-domain audio separation network; Conv- TasNet, an improved convolutional variant of TasNet; and Dual-Path Recurrent Neural Network (DPRNN), a model that addresses the limitations of traditional RNNs in handling long sequences. Experimental results using the WSJ0-2mix dataset are then presented, highlighting the strengths and weaknesses of each model. Finally, this paper introduces possible improvements to make the discussed models more efficient and applicable. The findings demonstrate improvements in isolating the target speaker's voice, making these models suitable for enhancing conference communications and other multi-speaker scenarios.