Handling Data Imbalance in Video Satisfaction Analysis
P.D. Thimira Madusanka, U. A. Piumi Ishanka · 2024
Video satisfaction analysis could be beneficial for content creators to improve their videos by using comments. Sentiment Analysis (SA) is used to ascertain the sentiment or emotional undertone of a text. In the realm of SA, data imbalance is one of the major problems in SA with Machine Learning (ML) models. This study focuses on comparing the performance of different Sampling Techniques used to address the data imbalance issue with the Sinhala language texts related to video satisfaction. Initially, the study created a labeled dataset with 3680 instances which was created by using YouTube video comments. In this experiment, the performance of seven different Sampling Techniques is investigated. SMOTE, RCSMOTE, ADASYN, Tomek Links, Condensed Nearest Neighbor (CNN) Sampling, NearMiss and SMOTE-Tomek sampling were used as Sampling Techniques for this study. These techniques are evaluated with six different ML models: Naϊve Bayes (NB), Support Vector Machine (SVM), Decision Trees (DT), Logistic Regression (LR), Random Forest (RF), and XGBoost. For individual Sampling Techniques, the Tomek Links undersampling technique performed the highest accuracy at 8 0. 7 0% in the Random Forest (RF) algorithm.