Summarization of Videos from Online Events Based on Multimodal Emotion Recognition

Amir Abdrahimov, Andrey V. Savchenko · 2022 International Russian Automation Conference (RusAutoCon) · 2022

In this paper, we propose a novel video summarization technique for automatic affect analysis of participants of an online event. At first, face verification neural network is used to cluster facial regions that correspond to each participant. Next, emotional features are extracted from each face by using EfficientNet model obtained in the previous paper of the author. The features of several consecutive frames are combined into a single descriptor that is used to classify emotions. In addition, audio features are extracted using wav2vec, and an ensemble of audio and video classifiers predicts emotions for each face. Finally, dependence of these emotions on time is visualized in special color charts. In the experimental study with the AFEW dataset it was demonstrated that the proposed approach makes it possible to obtain the best-known validation accuracy 67.88%. The models were optimized using OpenVINO and gained reasonable performance even if Nvidia GPUs are unavailable. The models and source code are publicly available at https://github.com/amirabdrahimov/multimodal-emotion-recognition.

Read the paper · More papers on PaperTik