Utilizing Ensemble Learning to enhance the detection of Malicious URLs in the Twitter dataset
Rakhi Arora, Rishi Gupta, Pradeep Singh Yadav · 2024
Due of Twitter's widespread appeal as a platform for sharing activities, entertainment, and information, spammers have taken notice and are using it to propagate false information and harass clients. Researchers are faced with the issue of locating user profiles and malicious URLs on Twitter so that prompt action can be taken. Several prior studies have tackled this issue and stopped spammers' actions on Twitter by utilizing various tactics. In this study, we create multiple models to detect content and categorise the URLs in it as spam or not-spam, employing several features, including hybrid, content based, and profile based. To build a spam detection, we first gather and categorise a sizable dataset from Twitter. Next, we take different features out of the gathered dataset and combine them to generate a set of rich features. In addition, we utilise various ensemble learning methodologies to construct prediction models. We conducted a thorough analysis of various methods using the gathered dataset, evaluating their performance in terms of f1-score, accuracy, recall and precision. The outcomes demonstrated that distinct sets of learning strategies employed have produced better results in the classification of tweet spam. Most of the time, the numbers for several performance metrics are higher than 90%. According to these findings, utilising user profile, content, and hybrid data while detecting suspicious URLs helps in the development of more accurate prediction models. This paper provides a current overview of prominent ensemble algorithms used in malicious URL detection. Notably, the k-NN technique, when applied in both bagging and random forest ensemble methods, demonstrates the highest accuracy.