Machine Learning Approaches for Sentiment Analysis on Balanced and Unbalanced Datasets
Ahmed M. ElMassry, Abdulla Alshamsi, Ahmed F. Abdulhameed, Nazar Zaki, Abdelkader Nasreddine Belkacem · 2024
Sentiment analysis, sometimes referred to as opinion mining, is essential for understanding public opinion and attitudes toward various social topics and trends. This study aims to explore the effectiveness of machine learning (ML) models, namely support vector machine (SVM), long short-term memory (LSTM), and bidirectional encoder representations from transformers (BERT), in analyzing a dataset obtained from Kaggle, which contains 37,000 user reviews on the Instagram Threads app. After initial data cleaning and preprocessing, the dataset was partitioned into ${7 0 \%}$ for training and ${3 0 \%}$ for testing. Subsequently, the training set was used to create three datasets: a balanced dataset and two unbalanced datasets, one featuring $90 \%$ positive instances and the other featuring ${9 0 \%}$ negative instances. Subsequently, these datasets were used to train the three machine learning models mentioned above, resulting in nine different models. Evaluation metrics, including accuracy, precision, recall, and F1 score, were applied to assess model performance. The finetuned BERT model on the balanced dataset outperformed all the other models with an accuracy of $86 \%$, precision of $85 \%$, recall of $87 \%$, and F1-score of ${8 6 \%}$. Furthermore, these findings underscore the effectiveness of diverse ML techniques, particularly transformers, and the crucial role of data balancing in optimizing sentiment analysis tasks.