Twitter (X) Spam Detection Using Natural Language Processing by Encoder Decoder Model

S. Asha, Mohan Madhan, Hari Krishnan P B, Sai Gokul Hariharan, B Dharaneesh, Sai Jeet K · 2024

Applications for social networking, such as Twitter, have grown in popularity in several fields, including politics, religion, economics, and entertainment. There is a lot of information available because of this popularity spike, some of which might be accurate while others might not be. By identifying irrelevant people and their material, spam detection is essential in resolving this problem. But up until recently, most research has concentrated on using activity detection and related technologies to collect user profile data. However, these techniques might not work as well if user profiles show temporal dependency or don't accurately represent the stuff they create. By concentrating on user profile data and content-based spam detection, this study seeks to address this problem. It presents three noteworthy additions. First, it uses cutting-edge natural language processing (NLP) techniques to create an extensive dataset with a wide variety of content-based attributes. Second, it analyzes this dataset using a hybrid machine learning model that combines deep learning and machine learning techniques. The practical value of this approach is highlighted by extensive simulations, which show that modeling both profile and content-generated data together works better than using individual techniques, with a combined spam detection accuracy of over 98 %. Finally, the paper presents a new approach based on logistic regression that is backed up by mathematical formulas. With the use of this technique, it is possible to evaluate the dataset and determine the likelihoods that legitimate users will differ from spammers. Future user categories can be predicted using mathematical results by varying the settings for each dataset. As a result, this method shows itself to be adaptable and efficient in identifying and classifying various user groups.

Read the paper · More papers on PaperTik