Multimodal Architecture for Emotion Prediction in Videos Using Ensemble Learning
V. Santhi, Puja Saha · 2022
Systems for the collection of user-generated videos and automatic analysis of these are expanding rapidly day by day in smart industries. In this chapter, we are going to predict emotions carried by the videos, for example, joy, sadness, etc. We begin by presenting a carefully designed dataset with manual annotations that was gathered from a major video-sharing website and can be used as a foundation for new study. This dataset provides a wide number of features, ranging from audio characteristics to images to high-level semantic properties. Convolution neural networks (CNNs) are used for emotion recognition in an image and Support Vector Machine (SVM) and Multi-Layer Perceptron (MLP) are being used to extract audio data. Both the emotion outputs from frames of video and audio are combined together by using ensemble learning. The primary focus is on the investigation of multimodal architecture for emotion transaction analysis in smart industries, which can predict emotions in a video.