Sentiment Analysis for Banglish Text using Machine Learning Approach
Musfique Ahmed, Fardin Hasan Siam, Neamul Islam Fahim, Md. Awinul Hoque Utsha, Md. Mahin Khan, Mohammad Nurul Huda · 2024
Sentiment analysis is a new area of study that has gained many achievements in the processing of Banglish, or Bengali-English, which is a type of language used for informal communication. Detecting sentiment in users’ comments is actually very important both for businesses and researchers as the number of online interactions increases exponentially. The machine learning models that we considered in this study for Banglish sentiment classification included Logistic Regression, Decision Tree, Random Forest, Support Vector Machine (SVM), Naive Bayes, and XGBoost. Although these models are quite good, there is another ensemble technique that is based on using Logistic Regression, Mini Batch K-Means, and Gradient Boosting. It utilizes the uniqueness of each model to enhance accuracy and to gain more insights into the subtleties of Banglish sentiment analysis. Our ensemble method does way better than any other tested model we tried. Through this study, it has been shown that ensemble techniques are very effective in sentiment analysis and this is the basis for the future research in the area of mixed-language context user sentiment understanding that is very important for the general consumer behavior analysis and for the extracting of information about the social trends.