Study of the Application of Logistic Regression and Naïve Bayes Algorithms for Automatic Classification of User Reviews in Bulgarian

Irena Valova, Gabriel Kanev, Tsvetelina Kaneva · 2024

Sentiment analysis (SA) is a critical task in natural language processing (NLP), aimed at evaluating and classifying user opinions in the form of text. This paper explores various methods for constructing sentiment lexicons, including manual, corpus-based, and dictionary-based approaches, each with its strengths and limitations. A specific focus is given to the challenges and methods of analyzing Bulgarian texts, highlighting the linguistic complexities that impact sentiment analysis in this language. The study includes an experiment using Logistic Regression (LR) and Naïve Bayes (NB) classifiers on a Bulgarian dataset, with vectorization techniques such as CountVectorizer and TF-IDF. The results are compared and examined and the combination of an algorithm and vectorization most suitable for this Bulgarian sentiment analysis is determined.

Read the paper · More papers on PaperTik