Optimizing Bengali Sentiment Analysis: A Comparison of Countvectorizer and TF-IDF with Machine Learning

Umme Ayman, Anzir Rahman Khan, Mehraj Hossain Mahi, Tamanna Akter, Zannatul Mawa Koli, Md. Hasan Imam Bijoy · 2025

In the realm of NLP, sentiment analysis is an indicator for identifying emotions in human language. Compared to the English language, research on low-resource languages like Bengali has not reached its full potential yet. This study evaluates the performance of CountVectorizer and Tf-Idf vectorizers to analyze sentiment in Bengali text accurately. Here we have applied SVM, MNB, RF and Xgb to the self-generated Bengali text dataset to evaluate the effectiveness of these feature extraction techniques. These techniques transform Bengali text into numerical representations for precise sentiment classification, addressing challenges like morphological complexity and syntactic nuances, highlighting the need for accurate representation of the text's sentimental features. Before applying machine learning algorithms, the dataset is preprocessed by utilizing several text preprocessing techniques to enhance its performance. We have obtained the highest accuracy of 91.15% with the MNB using the CountVectorizer strategy, outperforming the Tf-Idf slightly, which reached the accuracy of 91.00% is obtained with the MNB. The overall accuracy with CountVectorizer is 85.38%-91.15% and with Tf-Idf is 85.34% -91.00%. It is also observed that CountVectorizer has performed better than Tf-Idf comparatively. Eventually, the findings highlight the strengths and limitations of each technique, which are imperative for future research and study in NLP for Bengali. This work advances Bengali sentiment analysis and improves sentiment analysis's accuracy by determining the most effective feature extraction method.

Read the paper · More papers on PaperTik