Exploring Fraudulent Content on Social Media: Employing Unigram Precision and Machine Learning to Identify and Dissect Deception

B. Jeevashri, Uma Priyadarsini P. S · 2025

Online social media websites have become significant conduits for worldwide knowledge diffusion in the modern day. Some people misuse this opportunity by spreading false information to damage reputations and make money. This work proposes a novel hybrid approach for identifying false information on social media by integrating Term FrequencyInvers Document Frequency (TF-IDF) with N-grams and Word Vector (Word2Vec) for feature extraction. In contrast to previous methods that use only statistical or deep learning models, this paper proposes a hybrid feature extraction method that identifies statistical and semantic features to enhance detection performance and scalability. The preprocessed data is tested using six supervised machine learning (ML) models for content credibility evaluation. The preprocessed data is evaluated using six supervised machine learning (ML) models for content credibility assessment. The experimental results show that unigrams and the Random Forest (RF) model achieve the best accuracy (99.71 %), outperforming conventional fake news detection techniques. In contrast to current methods, this model efficiently utilizes statistical and semantic features, thus being more robust and scalable for detecting misinformation. These results show that the suggested method can significantly enhance the accuracy of fake news labeling, leading to more efficient and automated fact-checking systems.

Read the paper · More papers on PaperTik