Comparison of Classification Methods Using Feature Selection for Smartphone Sentiment Analysis
Ghiyalti Novilia, Ramzi Adriman, Taufik Fuadi Abidin · 2022
The determination of features is a major issue in the process of sentiment analysis classification. The right features can be chosen to reduce the dimensions of the dataset, making the classification stage more efficient and increasing the accuracy value. The study employed two methods for sentiment vector formation: first, N-Grams features yielded 6 and 18 features, respectively, and second, TF-IDF and Query Expansion Ranking (QER) yielded vectors of 10,20,30,40, and 50 results from word warfare. The f-measure value generated by Naive Bayes on 6 features is 0.73, while the f-measure value generated by SVM is 0.76. On 18 features, the f-measure value of Naive Bayes is 0.71 and 0.74 for SVM. The f-measure value in the combination of Naive Bayes and QER is highest in the 20k at 0.88, as well as on SVM and QER f-measure values of 0.91 for all k. In the n-grams feature, SVM accuracy on 6 and 18 features is higher than Naive Bayes, while QER has the highest accuracy value. In this study, the SVM algorithm outperformed Naive Bayes in the analysis of smartphone problem sentiment. When QER feature selection is used, accuracy improves.