Sentiment Analysis Optimization Using Vader Lexicon on Machine Learning Approach
Dhendra Marutho, Muljono Muljono, Supriadi Rustad, Purwanto Purwanto · 2022
The current increase in global internet users, equivalent to 59.5 percent of the world’s population, affects active internet users globally. Almost some users have expressed their opinions on the internet, including in text, in the form of positive or negative views. Sentiment analysis aims to polarize information in the form of text in the state of positive, negative, and even neutral opinions. Vader Lexicon can extract sentiment features properly on datasets without labels so that datasets can be directly classified using machine learning methods, including naïve Bayes, support vector machine (SVM), random forest, neural network, decision trees, k-nearest neighbor (k-NN). The selection of extraction features among TF-IDF, word2vec, and word embedding (one hot coding) also increases the accuracy of machine learning methods. The experimental results show that the feature extraction of TF-IDF with the SVM method is the highest. The evaluation of model performance using k-fold cross validation obtained 92.6% accuracy for unlabeled and 80.3% for labeled datasets.