Text Emotion Mining on Twitter (Preprint)

Suboh M. Alkhushayni · 2019

BACKGROUND Although the emotion model chosen for this study has several benefits in the context of this work, several research groups have proposed alternative sets of “basic” emotions that range in size from two to over 18 categories. Ortony and Turner construct a table of several of these emotion sets developed in the late nineteenth and twentieth centuries [5]. From this table, we can generate a frequency distribution, from which we conclude that the six most common “basic” emotions are: fear, anger, disgust, sadness, joy, and surprise. Note that these six most popular emotions are a direct representation of the model of basic emotions proposed in 1972 by P. Ekman, W. Friesen, and P. Ellsworth [2]. Shahraki and Zäiane, creators of the CBET (cleaned balanced emotional tweet) dataset, also recognize that basic human emotions have been a controversial issue among scientific studies and admit that many models of the human emotional spectrum exist [6]. Like Shahraki and Zäiane’s work on the CBET dataset, we base our emotion set on Ekman’s basic emotion model (anger, disgust, fear, joy, sadness, surprise) in this work due to its clear distinction between emotions and relative simplicity of the model [6]. Like [6] guilt was included as a basic emotion since guilt is another commonly recognized basic emotion by many psychologists such as C.E. Izard [3]. Since the far-reaching goal of this research includes assisting mental health professionals with the detection of those who are suffering or may suffer from symptoms of depression, guilt was included to pinpoint tweets that might help alert psychologists detect such conditions. Due to the complexity of multi-class supervised machine learning classification, few works have been done with the focus of such classification. Problems can arise such as needing large amounts of training data, computing power, and time. Within these works, even fewer have been done in the specific context of Twitter. Most notable among these is the work of Shahraki and Zäiane on the CBET dataset [6]. Their work consisted of testing and evaluating several lexical and learning-based methods on the CBET dataset, which included classification using both Naïve Bayes and binary SVM classifiers. Classification problems such as this are often addressed using one of two primary approaches. The first of these is a lexical approach, which employs a vocabulary to tag a piece of text with a corresponding emotion [6]. Also available is a learning approach, which uses a (set of) trained machine learning algorithm(s) to predict the emotion of test data based on patterns extracted from training data [6]. The classification method employed on our dataset is a hybrid of the method used by Kinsler [12] in that the method is based on ensemble of multi-class classifiers. The “votes” in this case are the tags outputted by several machine learning algorithms, from which the most popular is selected as the tag for the inputted tweet. We say that our system is primarily based on Kinsler’s, but there exists a key difference between the application methodologies. Kinsler uses his voting classifier for sentiment analysis, wherein binary (or n-class where n = 2) classification is used [12]. Given the set of emotions used to construct our dataset, we require a seven-class classification system, which significantly increases the problem’s complexity. OBJECTIVE Emotion mining, or emotion identification, generally describes the practice of determining and analyzing the feeling(s) expressed towards companies, products, events, people, etc. [1] Emotion mining can be used to label a variety of forms of human input, including spoken words, written words, and facial expressions with one or more classifications from a set of basic emotions. According to Neel Burton of Green-Templeton College, these basic emotions are “‘hardwired’, where each basic emotion corresponds to a distinct and dedicated neurological circuit” [4]. However, the disagreements on a set of primitive or basic emotions, as well as how one would classify an emotion as basic, is debated to this day among emotional psychologists. Thus, choosing the “proper” set of emotions typically incorporates subjectivity. This study addressed the written word and classification of human writing based on a unique set of seven basic emotions consistent with the model proposed by Ekman [2] with the addition of guilt. METHODS Creating the Emotionally Tagged Twitter Dataset (ETTD) Emotion Classification Lexical Approach Learning-Based Approach RESULTS At the time of this study, the most modern publicly available dataset related to emotion classification was the CBET dataset [6]. This dataset, containing 76,860 tagged tweets, was created with an emotion set differing from the set used in this study only with the inclusion of love and thankfulness as emotion classes. The results of binary SVM classification on the CBET dataset after five independent trials were a precision score of 45.01%, recall score of 42.26%, and F1 score of 43.59%. Again, five independent trails of the lexical-based method described in this work were conducted on the common emotion classes between CBET and ETTD (Anger, Fear, Guilt, Sadness, Disgust, Joy, and Surprise). From this, the values (A = 35.57%, P = 42.5%, R = 34.65%, F1 = 35.1%) were attained. Next, an ensemble classifier was trained and tested on the same ratio of the dataset as used in this work (75% for training to 25% for testing). Again, five independent trials were conducted, from which the values (A = 45.67%, P = 48.82, R = 45.09, F1 = 46.03) were attained. Due to the similar construction of the CBET dataset and ETTD, the similar results were as was expected from these trials. Also, due to the similarity of these datasets, confusion matrices were generated for this experiment, as shown in tables 6a and 6b. CONCLUSIONS Determining the emotion in a piece of social media text is critical since it is commonplace for a large portion of the global population to communicate their thoughts and feelings on such media in response to daily events. The concept of emotion mining, though young, is a fascinating challenge which incorporates a plethora of professional disciplines beyond computer science. This study addressed this challenge with the ETTD, a corpus of 24,500 tweets labeled with one of seven emotional tags: joy, fear, surprise, anger, sadness, disgust, and guilt. Next, a lexical approach, in which we use weighted word-emotion association vectors, was tested as a preliminary attempt to accurately and reliably classify tweets. After this, a supervised machine learning-based approach to the problem was assessed. The algorithms used in this approach were trained on 75% of the ETTD to extract patterns from the training data and then use such patterns to classify new tweets. Each of the methods showed promising results on the ETTD and achieved similar results when applied to the CBET dataset and the dataset created by [8]. Significant overlap of guilt with sadness, anger, and disgust is shown in the confusion matrices produced by this work, which may suggest that guilt is in fact a combination of such emotions.

Read the paper · More papers on PaperTik