Emo-BERT: A Multi-Modal Teacher Speech Emotion Recognition Method
Gang Lian Zhao, Yinan Zhang, Jie Chu · 2023
In the field of education, the existing automatic teacher emotion recognition is mainly realized by unimodal information such as audio, text, and facial expression. The effectiveness of multimodal speech emotion recognition methods has been demonstrated in the industrial field. To construct an accurate and feasible multimodal speech emotion recognition method, a neural network named Emo-BERT, which combines audio features with text content for teacher speech emotion recognition is proposed. The proposed Emo-BERT can deeply integrate the audio and text features of teacher utterance, and realize more efficient teacher speech emotion recognition. The audio encoder in Emo-BERT can effectively integrate the frame-level features and utterance-level features of teacher utterance through the attention pooling block, so that Emo-BERT can select valuable features more effectively. The proposed Emo-Bert achieves unweighted accuracy of 77.6%and 64.0% on IEMOCAP and MELD. A teacher emotion dataset (TED) was established from real teaching scenes to evaluate proposed method. The proposed method fills the gap of multimodal teacher speech emotion recognition, our proposed method has higher accuracy.