Speech Separation in Time-Domain Using Auto-Encoder

Jaipreet Kour Wazir, Pawan Kumar, Javaid Ahmad Sheikh, Karan Nathwani · 2024

One of the primary challenges for automatic speech recognition (ASR) and speaker recognition is speech separation, which tracks and identifies a particular speaker's speech while multiple speakers speak simultaneously. Deep learning methods have shown excellent results in speech separation processing. Auto-encoders have made significant advances in deep learning. They outperform recurrent and convolutional models in multiple tasks and utilize parallel processing. This paper presents the auto-encoder in the time domain for speech separation. Experiments demonstrate that the proposed approach performs at the state-of-the-art (SOTA) level on the publicly available TIMIT data set. Objective and subjective parameters have been calculated for the proposed approach.

Read the paper · More papers on PaperTik