Speech Separation and Emotion Recognition for Multi-speaker Scenarios
Rong Jin, Mijit Ablimit, Askar Hamdulla · 2022
Emotion recognition in multi-speaker scenarios has been a hot topic of continuous research. In the single-channel overlapped speeches, the recognition accuracy is greatly degraded compared to clean speeches. Therefore, a speech separation task for multi-speaker separation at the front end of emotion recognition is necessary. The overall efficiency is the combined effect of separating a single speech from overlapped voices and emotional recognition. The speech separation experiments in this paper are based on Conv-TasNet, using the WSJ0-2mix dataset as training data to obtain the trained separation model; meanwhile, the speech emotion recognition experiments are inherited from Wav2vec-2.0 pre-trained model with a multi-task learning framework to recognize both text and emotion. Through the experiments, the possibility of bridging the two tasks and the usability of the separation and recognition results for realistic scenarios are demonstrated, and subsequent work can be carried out on improving the robustness and application to other realistic scenarios.