A Machine Learning Approach to Spam Detection in Social Media Feeds
Anindya Sudhir, L. E. Joseph, N. Ahmad, S. Awasthi, Vivek Agarwal, S. P. Raja, Zoran Stamenković · 2023
Twitter, Facebook, and Reddit are popular social networking sites for online interaction. These platforms use spam filtering techniques to identify and block spammers. However, discovering spammers is challenging in Online Social Networks (OSNs). The most commonly used technique to identify spammers is classification-based supervised techniques. However, these sys- tems are subject to "data manipulation," "imbalanced datasets," "data marking," and "spam flow" restrictions. Malicious users exploit implicit trust ties between users to spread false infor- mation, post or tweet malicious websites, and contact legitimate people without their consent. Using unlabeled URLs in a semi- supervised manner and synthetic harmful URLs produced via adversarial learning can reduce the need for labelled samples and solve the imbalance problem. The characteristics of spam users on Twitter are investigated to enhance the effectiveness of current spam detection systems. A collection of new features, more reliable and effective than previously used ones, are used to identify Twitter spammers. The suggested attributes are evaluated using well-known machine learning classification algorithms, such as Support Vector Machines (SVM), several Naive Bayesian (NB) methods, K-Nearest Neighbor (KNN), and Random Forest (RF).