Comparison of K-Nearest Neighbor and Logistic Regression Algorithms on Sentiment Analysis of Covid-19 Vaccination on Twitter with Vader And Textblob Labeling

Fadhilah Fazrin, Oktariani Nurul Pratiwi, Rachmadita Andreswari · 2022

In this analysis, the methods used are the K-Nearest Neighbor classification method and the Logistic Regression classification method with data taken on the twitter application. This study examines the level of accuracy in public sentiment regarding covid-19 vaccination with positive and negative labels. The AUC value in the KNN algorithm with TextBlob labeling is 0.765 with and 0.76S for VaderSentiment labeling are both included in the fair classification criteria. Meanwhile, the Logistic Regression algorithm produces an accuracy of 84.97% with a ratio of 90:10 for Labeling TextBlob, while for Labeling VaderSentiment with a ratio of 90:10 results in an accuracy of 85.22%. Both algorithms are validated using K-Fold Cross Validation with a fold count of 10. The comparison results obtained when conducting an evaluation with the confusion matrix showed that the Logistic Regression algorithm with VaderSentiment labeling had the highest accuracy value compared to the K-Nearest Neighbor algorithm with TextBlob and VaderSentiment labeling.

Read the paper · More papers on PaperTik