Automating Video Frame Analysis for Emotion Recognition and Captioning in Real Time
Rahul Jiandani, Vanshika Nijhawan, Archana Lakhe · 2024
Facial Emotion Recognition (FER) is an evolving field within machine learning and computer vision, with applications in personalized advertising, customer satisfaction analysis, and real-time feedback in interactive systems. This paper explores recent advancements in automatic FER, emphasizing the impact of deep learning architectures on performance. We compare the effectiveness of traditional machine learning models, like Decision Trees (39.44%), Random Forests (56.77%), and SVMs (58.93%), with modern deep learning approaches. Although deep learning models show improved accuracy, they continue to face challenges, particularly overfitting, as seen in the gap between training (84.732%) and validation accuracy (63.250%). Our work develops an automated system for real-time FER that improves emotion detection accuracy and real-time processing efficiency. The system analyzes video frames to recognize emotions, generates captions, and identifies the most impactful video segments for eliciting emotional responses. This approach will refine customer experiences and revolutionize marketing strategies through personalized and immediate interventions. To improve model generalization and address challenges in cross-cultural and spontaneous emotion recognition, we leverage the Indian Spontaneous Expression Database (ISED), which provides diverse, high-resolution videos of spontaneous facial expressions but presents challenges like imbalanced classes and subtle emotional nuances. Advanced techniques, including attention mechanisms, hybrid models combining CNNs and RNNs, data augmentation, and Bayesian Optimization, are employed to bridge the gap between training and validation accuracy. This improves validation accuracy from 60-64% to closer to training accuracy (~80-82%), effectively overcoming the limitation of the dataset and making the model more robust.