Research on Multi-person Speech Recognition Based on Deep Learning

Xueying Li, Changliu Niu, Xinyue Wang, Xi Jiang Han · 2023

We present a system for speech separation and speaker recognition, a recognition technology for multi-person mixed speech signals. It’s mainly used for identification of multi-person speech signals, however, it does not pay attention to the recognition of speech content. The proposed system is divided into two parts: speech separation and speaker recognition. For the first part of the task, we propose a Convolutional Neural Network-Gated Recurrent Unit-Attention (CNN-GRU-Attention) model that uses a convolutional neural network to convert the audio signal into a logarithmic spectrogram as input and a GRU to model the timing information. The presented method delivers better generalization ability and speech separation effect. As for the second part of the task, To address the problem that the convolution process generates a large number of channels containing redundant information such as noise and silent segments, the attention mechanism module SENet is introduced to improve the model, giving more weight to the channels containing important information and improving the recognition effect.

Read the paper · More papers on PaperTik