Comparative Analysis of Text based Emotion Detection on GoEmotions Dataset
Vinod Kumar, Akshit Bansal, Umang Gupta · 2023
Today’s environment relies heavily on social media, blogs with customer reviews, tweets, and comments for important communication. Understanding the feelings conveyed through these tweets and comments is crucial for improving a variety of fields, such as the detection of harmful online behavior (hate speech and the use of abusive language), a better understanding of customer review content, the development of chatbots with empathy, and many others. As a result, we introduce Go Emotions, the biggest humanly annotated dataset of 58k English Reddit comments that have been classified as either neutral or one of 27 emotion categories. The GoEmotions dataset’s distinguishing feature is its emphasis on capturing a broad spectrum of emotions. GoEmotions includes a more varied and nuanced set of 27 emotion categories, in contrast to many other emotion datasets that often concentrate on few basic emotions (such as happy, sadness, rage, etc.). Not only do these categories cover fundamental emotions, but also more subtle ones like appreciation, amusement, gratitude, love, and others. Due to a wide range of categories, on the GoEmotions dataset for emotion classification, basic models like SVM (Support Vector Machines) and Naive Bayes might do worse than advanced models like CNNs, RNNs, and Transformers. Thus, applying the methodology from the base paper (BERT) and other advanced models, we present a comparative analysis of the same.