Fake Detection for Tweets and News with Different Domain Datasets
Sehrinaz Koca, Ilyas Cicekli · 2024
Due to the widespread application and usage of social media, the issue of whether shared data contains genuine or misleading information has surfaced. Fake detection applications are utilized to identify fabricated Twitter(X) posts and deceptive news articles related to politics or natural disasters. These applications rely on algorithms designed to categorize text into two classes and the categorization process is facilitated by models trained with a designated training dataset. Various machine learning algorithms are employed within fake detection applications. Typically, trained models are evaluated using a test dataset within the same domain as the training data. However, sourcing a suitable training dataset for every fake detection application may not always be feasible. This paper aims to assess the efficacy of different machine learning algorithms by employing training datasets from different domains and comparing their performance across different test datasets. Six distinct algorithms are applied to four different datasets to evaluate their performance in fake detection classification tasks. Performance metrics are thoroughly analyzed concerning both datasets and algorithms. The datasets encompass diverse subjects such as Covid19, politics, economy, and tornadoes. Although classical machine learning algorithms do not perform very well when train datasets come form different domain, deep learning based machine learning algorithms produce good results with the train datasets from different domains.