Comparative Analysis of Machine Learning Algorithms for Arabic Sentiment Analysis on Imbalanced Social Media Data

Basheer Almuhaya, Bishal Saha, Manbir Kaur, Mahmood A. Bazel, Rehab Ghaled Mohammed · 2024

Sentiment analysis is crucial in understanding public opinions and attitudes on social media platforms. However, dealing with imbalanced datasets, especially in the context of Arabic sentiment analysis, poses significant challenges. In this study, we utilize five well-known machine learning algorithms: support vector machine, logistic regression, random forest, multinomial naive Bayes, and Knearest neighbors, in combination with appropriate preprocessing techniques and pipelines, such as text cleaning, normalization, tokenization, and feature extraction using TFIDF vectorizer. Additionally, we addressed the imbalanced data issue through synthetic minority over-sampling technique and random under-sampling. The evaluation was conducted on the SS2030 Twitter dataset, consisting about 2,436 and 1,816 positive and negative tweets in Arabic. This study provides insights into the performance of various algorithms and the effectiveness of sampling techniques in mitigating the challenges posed by imbalanced data in Arabic sentiment analysis.

Read the paper · More papers on PaperTik