Video-Audio Emotion Recognition Based on Feature Fusion Deep Learning Method

Yanan Song, Yuanyang Cai, Lizhe Tan · 2021

In this paper, we propose a video-audio based emotion recognition system in order to improve the successive classification rate. The features from audio frames are extracted using Mel frequency Cepstral coefficients (MFCC) while the features from video frames are extracted from VGG16 with pre-trained weights on the ImageNet dataset [17]. Then recurrent neural networks (RNN) are further applied to process the sequence information. The outputs of both RNN are fused into a concatenate layer and then the final classification result is obtained by the softmax layer. Our proposed system achieves 90% accuracy based on the RAVDESS dataset for eight emotion classes.

Read the paper · More papers on PaperTik