Analysis of Feature Extraction by Convolutional Neural Network for Speech Emotion Recognition

Daisuke Horii, Akinori Ito, Takashi Nose · 2021 IEEE 10th Global Conference on Consumer Electronics (GCCE) · 2021

In recent years, neural network-based methods have become the mainstream for speech emotion recognition, and Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN) are often used. In particular, the use of CNN has become popular in recent years. Many researchers are trying to extract better features using CNN for spectrograms or MFCCs, but it is not verified what kind of features CNN extract and whether they are superior to conventional features, F0, energy and MFCC. In this paper, we performed emotion recognition using two models, one with CNN and one without, and then analyzed the features extracted by CNN.

Read the paper · More papers on PaperTik