Comparison of Feature Selection for Naive Bayes Classification Method in A Case Study of The Corona virus Lockdown

Adian Fatchur Rochim, Ratri Kusumastuti, Ike Pertiwi Windasari · 2021

Classification in sentiment analysis often involves less relevant features for the modeling process. This causes the accuracy obtained to be not optimal. Therefore, a feature selection method is needed to sort out features that have high relevance to the dataset. This research aims to compare the accuracy between three different methods. They are Naive Bayes Classification without using any feature selection, using Information Gain, and using Chi-Square feature selection. The datasets used are sentiments related to the lockdown as a policy for the Coronavirus pandemic from Twitter. Feature selection methods affected the accuracy by filtering features and sorting the most relevant features based on its algorithm. The results showed that the average accuracy of Naive Bayes without feature selection, using Information Gain, using Chi-Square based on both Indonesian and English datasets were 63.2 %, 64.2%, and 65%, respectively.

Read the paper · More papers on PaperTik