Real-Time Twitter Corpus Labelling Using Automatic Clustering Approach
Itisha Gupta, Nisheeth Joshi · International Journal of Computing and Digital Systems · 2021
In this paper, we present a novel automatic labelling approach for the classification of large amount of unlabelled real-time twitter datasets for textual-based Twitter Sentiment Analysis.The tweets are labelled or classified as Positive, Negative or Neutral using the novel automatic approach.The proposed approach applies an unsupervised clustering technique that would generate clusters based on the underlying patterns (finding similarities between tweets) in the collected twitter corpus.Twitter search API is used to collect real-time English tweets on several topics such as "#Demonetization", "#lockdown", and "#9pm9minutes" by the use of search operator.To analyse the sentiment from real-time tweets, labelling of the corpus is required.Manual annotation of large twitter corpus is time and labor-intensive.Moreover, domain experts are needed for labelling of tweets belonging to a particular domain.Thus, in this work, we propose the use of the improved K-mean clustering approach, which is an unsupervised way of labelling corpus, which could then be used for learning supervised models such as SVM for sentiment analysis.To make the corpus ready for clustering and to get quality clusters, we have applied some basic to advanced cleaning operations known as tweet normalization.Furthermore, extensive feature engineering is conducted to generate different types of features including POS-based (Part-of-Speech), ngrams, Twitter-specific, negation, and lexicon-based features from our collected unlabelled twitter corpus.Those features act as input to the K-mean clustering algorithm and help it in identifying patterns from the data for cluster generation.Moreover, we handle an important linguistic phenomenon namely negation before the cluster generation.Our main focus is on handling those negation tweets in which negation presence has literally no sense of negation (negation exception cases).At the end, cluster analysis is done manually to find out the sentiments expressing by tweets in a particular cluster.Accordingly, cluster classification is done and each cluster is assigned one class that is Positive, Negative, or Neutral.The main contribution of this work is the idea of amalgamation of extensive feature engineering and negation modelling with the unsupervised K-mean clustering approach for classification of large unlabelled twitter corpus.A comparative analysis of our proposed approach is done with or without negation exception cases and with random K-mean using only conventional TF-IDF as features.The proposed automatic labelling approach produces substantial results in terms of cluster quality assessed through two evaluation metrics known as inertia and silhouette score. .