An Empirical Study on the Impact of SMOTE on Imbalanced Text Features for Review Rating Prediction in Bengali
Farhad Uz Zaman, Md. Zahid Hossain, Md. Kamrozzaman Bhuiyan · 2023
The sole method for offering feedback regarding a product is by means of reviews. Reviews along with their ratings serve as a means for customers to inform others about the product’s quality, while manufacturers can use these reviews to improve their products and grow their businesses. When a new consumer does not have time to read all of the reviews supplied by other customers in order to evaluate a product’s quality, they frequently rely their purchase decision solely on the numerical rating. Manufacturers also prioritize numerical ratings over written reviews for business evaluations. In this paper, we propose a machine learning model that has been trained to predict star ratings from Bengali customer reviews. The study used a dataset collected from Daraz1, one of Bangladesh’s top online retailers. We pursued two distinct approaches: one, the Baseline method utilizing imbalanced data, while the second technique incorporated the Synthetic Minority Over-sampling Technique (SMOTE). We conducted experiments employing four machine learning algorithms: Multinomial Naive Bayes, Logistic Regression, Random Forest, and Support Vector Machine. Notably, the Support Vector Machine model in the SMOTE approach outperformed all the other techniques achieving an accuracy of 92.80% and an F1-Score of 92.67%.