Speech Emotion Recognition Based on Linguistic Features of Spoken Language
Keita Kishi, Kiyohide Sato, Tetsuo Kosaka · 2024
We study speech emotion recognition based on linguistic features that consider the spoken language in Japanese. In this approach, speech recognition is used to convert speech into text. The text obtained was then used for emotion recognition. Japanese is expressed differently in written and spoken forms. It is important to consider the differences for speech emotion recognition when using linguistic features. In this study, we used BERT, a deep learning model, for speech emotion recognition. Because emotional speech is expressed in spoken language, the pre-training model of BERT should also be trained in spoken language. However, conventional methods typically use a pre-trained model trained on a written language. Therefore, we attempted to improve recognition accuracy by using a pre-training model that is trained using text on SNS, which is considered to be similar to spoken language. The results indicated that the model pre-trained with data from SNS achieved a recognition rate of 76.75%, whereas the model pre-trained with conventional written language achieved a recognition rate of 62.75%, showing a performance improvement. We also evaluated this method using an open task and demonstrated its effectiveness.