A Model for Song Recommendation Based on Facial Emotion Analysis and Musical Emotion
International journal of intelligent engineering and systems · 2024
Enhancing user experience in music listening through facial emotion recognition presents a novel avenue for curating personalized song playlists, obviating the need for traditional search methodologies like keyword inputs or manual playlist exploration.This research pivots towards leveraging advanced pre-trained models, specifically ResNet18 and VGG19, for the accurate detection and analysis of listeners' emotions via facial expressions.By integrating these models with the Facial Emotion Recognition (FER) 2013, and The Extended Cohn-Kanade Dataset (CK+) dataset, our study aims to precisely identify a range of emotions including happiness, anger, sadness, surprise, disgust, and neutrality.In parallel, we employ Support Vector Machine (SVM) and Linear Discriminant Analysis (LDA) techniques to discern musical emotions from Spotify's dataset, encompassing moods such as happiness, sadness, energy, and calmness.A unique set of rules is formulated to amalgamate the insights from facial emotion recognition with musical emotion data, facilitating the automated suggestion of playlists that resonate with the user's current emotional state.Our experiments demonstrate notable results with the ResNet18 model achieved 72.67% accuracy and a 72.30% F1 score on the FER2013 dataset, with significant improvements on the CK+ dataset, reaching 96.97% accuracy and a 96.60% F1 score.The VGG19 model similarly excelled, particularly on the CK+ dataset with 95.96% accuracy and a 96.60% F1 score.On the FER2013 dataset, the VGG19 model achieved an accuracy of 72.64%, and an F1 score of 72.60%.Our music emotion recognition approach, using SVM on Spotify's dataset, achieved an 86.96% accuracy and an 84.54% F1 score.When using the LDA model on the same dataset, the accuracy was 79.00% and an F1 score of 76.72%.To bring this concept to fruition, we have developed an intuitive web application that allows users to capture their facial expressions via their device's camera, subsequently offering song recommendations tailored to their detected mood.