A Comparison of Text Classification Methods: Towards Fake News Detection for Indonesian Websites

Trisna Ari Roshinta, Hartatik Hartatik, Elya Kumala Fauziyah, Ivan Fausta Dinata, Nurul Firdaus, Fiddin Yusfida A’la · 2022

Fake news reports false or distorted information that aims to mislead us and undoubtedly has a negative impact on society. For example, medical research in Taiwan shows that fake news about the COVID-19 vaccine reduces the number of doses absorbed by the public significantly as those exposed to fake news become hesitant and even anti-vaccine. Therefore, early automatic detection of fake news on the internet is crucial. In detecting fake news, machine learning algorithms, especially text classification algorithms, are used as a solution for this problem. Searching for the best method with high accuracy needs to be done continuously. This paper compares four main algorithms, namely Support Vector Machine (SVM), Stochastic Gradient Descent (SGD), Logistic Regression (LR), and Naïve Bayes. The experiment was carried out using 200 Indonesian news datasets, consisting of 100 fake news and 100 real news. The performance matrix of each algorithm was evaluated with accuracy, recall, precision, and F1-score as the harmonic mean. The results showed that Logistic Regression was able to separate fake news and real news with the highest F1-scorc reaching 90.9%. This paper also proposes a framework of detecting fake news that can be implemented on public websites.

Read the paper · More papers on PaperTik