Misleading multimodal news dataset for detecting fraudulent content
Deepika Varshney, Mohit Aggarwal, Neetu Saradana · 2024
With the rise in the popularity of social media, misinformation in the form of false or misleading content such as hoaxes, conspiracy theories, fabricated reports, clickbait headlines, and even satire is being widely disseminated with the goal of improving social media marketing effectiveness, boosting online traffic, increasing the number of followers for a page or business, creating a distraction, eliciting an emotional response, and shaping or changing public opinion. One of the greatest hurdles to their identification is the dearth of comprehensive datasets on misleading video and fake image detection that incorporate multiple categories. Hence, using a data extraction tool, a dataset consisting of forged images and misleading videos was created. The dataset consists of images, videos, and text from different social media platforms, thereby leading to multi-modality. The dataset created could now help in the identification of misleading information existing in any type of communication. This chapter presents the dataset along with the evaluation of the dataset to show where the dataset stands in terms of training and testing a model for the identification of fake news or misleading information. It also proposes the novelty of the dataset by testing variety of models and algorithms to predict misleading information. The algorithms used by different researchers and scholars have been taken into consideration and the results obtained are quite promising. Of all the models and algorithms applied, random forest and decision tree acquired the highest accuracy of 99.39%.