Classification of Twitter Contents using Chi-Square and K-Nearest Neighbour Algorithm

Yunita Dwi Setiyaningrum, Annisa Fitrianingtiyas Herdajanti, Catur Supriyanto, Muljono Muljono · 2019 International Seminar on Application for Technology of Information and Communication (iSemantic) · 2019

A survey conducted by Hootsuite in 2019, that there are 3,484 billion active social media users. The percentage of active social media users has increased by 15% from 2018. One way or action is to anticipate negative things that quickly circulate among the public such as fake news (hoax), reduced ethics and courtesy or referred to as negative content. From various types of social media such as; Facebook, Twitter, Instagram, and others, one of which Twitter has a percentage of 52% is a social media network service. In this study we used a dataset obtained from UCI in 2014 entitled “Twitter dataset for Arabic Sentiment Analysis”[7]. Category of twitter content using the Chi-Square algorithm and KNN as the method. The chi-square method is an algorithm for selecting features, and for the classification process the author uses the KNN algorithm. This study concluded that, the KNN algorithm with feature selection (chi-square) obtained the results of accuracy seen from the value of the nearest neighbor distance K = 3 obtained the same results with an accuracy value of 65.00%, which states that both methods with feature selection and without selection the feature in the dataset turns out that the method with feature selection has the advantage of being more efficient in the calculation process.

Read the paper · More papers on PaperTik