COMPARISON OF CLASSIFIERS FOR THE TASK OF TEXT TONALITY ANALYSIS
Nataliya I. Boyko, А-М. R. Kulinchenko, Kateryna Hazdiuk · Computer Science and Applied Mathematics · 2023
The study aims to determine the most effective classifier for the task of analyzing the tonality of the text. Naive Bayes, logistic regression, decision trees, and random forests are among the ones specified in work for comparison. Sentiment analysis was chosen for the task of analyzing the tonality of the text. A set of movie reviews provided by IMDB critics was selected as the basis for the research. The objects of the study are chosen directly as classifiers, and the subject, accordingly, is the determination of their effectiveness when applied to the problem mentioned above. This chapter aims to introduce and evaluate classification methods in the context of sentiment analysis. Classifiers such as naive Bayes, logistic regression, decision tree and random forest were compared. Let’s move on to a more detailed description of each of them. This study determines the most effective classification algorithm for the analysis of the text’s tonality. This, in turn, enables programs that perform such analysis to improve the quality of text distribution into different groups. In this work, the classifier was determined which among the selected ones is the most effective – it is the method of logistic regression. During the implementation of this work, analyses of the relevance of the problem and scientific sources were conducted, among which the accuracy of the naive Bayes, logistic regression, decision trees, and random forest classifiers was investigated. Before direct classification, the distribution of the data set into training and test samples is given. Each classifier is trained with specific hyperparameters. Detailed analysis and data preparation for the binary classification task was also performed. In parallel, training classifiers and conducting experiments with each were carried out. The study results were discussed using complete statistics of all metrics and all selected classifiers. To improve the accuracy of classifiers, choosing appropriate hyperparameters for each type is necessary. Analysis of the words themselves is carried out. Statistical calculations of the terms used in positive and negative reviews were carried out, and accordingly, “word clouds” with the most used words were constructed. For a more detailed analysis, inconsistency matrices were also created for each method.