Research on Speech Separation Algorithms Based on Deep Learning

Beidan Liu, Zhiyong Zhang, Yuanlong Li, Jiaxuan Ran · 2025

This paper introduces an efficient speech separation algorithm based on optimized Long Short-Term Memory (LSTM) networks. The enhanced LSTM architecture improves accuracy and processing efficiency, enabling real-time applications for multi-speaker separation. The mixed speech signal is first processed via Short-Time Fourier Transform (STFT) for time-frequency representation. A 2D convolutional block extracts features, followed by a deep neural network for advanced processing, with feature fusion through a fully connected layer. The network-processed data undergoes an XOR operation with raw data, and 2D transposed convolutional operations reconstruct isolated speech signals. Experiments demonstrate high scale-invariant signal-to-noise ratio (SISNR) values: 10.27 dB (LSTM), 9.56 dB (BLSTM), and 7.32 dB (Transformer), validating the algorithm's efficacy in multispeaker separation.

Read the paper · More papers on PaperTik