Speech Emotion Recognition Using Speech Feature and Word Embedding

Bagus Tris Atmaja · 2019

Emotion recognition can be performed automati-cally from many modalities. This paper presents categoricalspeech emotion recognition using speech feature and wordembedding. Text features can be combined with speech features toimprove emotion recognition accuracy, and both features can beobtained from speech via automatic speech recognition. Here, weuse speech segments of an utterance where the acoustic featureis extracted for speech emotion recognition. Word embeddingis used as an input feature for text emotion recognition anda combination of both features is proposed for performanceimprovement purpose. Two unidirectional LSTM layers are usedfor text and fully connected layers are applied for acousticemotion recognition. Both networks then are merged to produceone of four predicted emotion categories by fully connectednetworks. The result shows the combination of speech and textachieve higher accuracy i.e. 75.49% compared to speech only with71.34% or text only emotion recognition with 66.09%. This resultalso outperforms the previously proposed methods by othersusing the same dataset on the same and/or similar modalities.

Read the paper · More papers on PaperTik