Twitter Sentiment Analysis, 2-Way Classification: Offensive or Not-Offensive?
Mestan Fırat Çeliktuğ · 2024
In the age of rapid information exchange, understanding public sentiment on social media platforms like Twitter has become essential for various stakeholders. This paper presents a comprehensive analysis of Twitter Offensive Language Detection problem using several robust, powerful and expressive machine learning models from different model families, including Artificial Neural Networks (ANN), Support Vector Machines (SVM), Decision Trees(DT), and Random Forests (RT). The primary objective is to evaluate the effectiveness of each model in classifying tweets as "offensive" or "not-offensive", which has a priceless precaution value for a great many institutions, corporations, and regulatory bodies with priority against Hate Speech Detection -as it is an aggravated form of Offensive Language Detection-. The studied annotated Twitter dataset is majorly utilized in the literature for a classification task including hate speech detection, which is different from our approach prioritizing the detection of all forms of offensive language regardless of hate speech.Upon tweet preprocessing that involves data cleaning (removal of several not-relevant patterns and stopwords), applying tokenization and stemmization ( Porter stemmization), the feature engineering is held in a way that the effects of unigram and bigram features, sentimental polarity-based feature extraction on overall and class-based accuracies are investigated, concluding with top-1500 unigram feature extraction for prevention of curse of dimensionality.Upon proper and well-administered hyperparameter tuning (incl. Early Stopping Method, Grid Search, Halving Grid Search) with the consideration of computation and storage resources, three (3) models (SVM, ANN, DT) reach 96% overall accuracy ratio (except RF with 95%). Overall, the better-grasped class during the learning process is the "Offensive" class (at best, DT with 97% "Offensive" class accuracy), in accordance with the overall class imbalance. However, worthy to note that, SVM outperforms other models (by 4% or 5%) with 94% "Not-offensive" class accuracy.One of the main contributions of the study is, for Twitter Offensive Language Detection, to demonstrate the power of proper, realistic hyperparameter tuning in the existence of sufficiently expressive model family by the same top-notch overall accuracy result of the three robust model families and 1% less overall accuracy of another robust model family.The class-based performance differences suggest an Ensemble approach combining the different robust model families might gain 1% or 2% overall accuracy based on the early literature.The work is made available on GitHub1for ease of use and access.