Robust Twitter Spam Detection Through Ensemble Learning and Optimal Feature Selection
Dharmesh Dhabliya, C. Karthikeyan, Gourav Sood, B Sivadharshini, Ayaan Faiz, Manita D Shah · 2024
Specifically, it delves into various ensemble learning techniques utilized for enhancing choosing features in the context of detecting junk on Instagram. Choosing the right features A further crucial stage in the process is choosing features that are both extremely useful and minimally associated with each other in order to enhance the categorization rate. This work provides a summary of various methods for selecting features, including filter, wrapper, and embedded approaches, along with unsupervised approaches. Features can be ranked through various approaches to choice such as chi-square, information gain, and random forest. This section delves into the concept of combining multiple feature selectors to enhance the reliability of outcomes and improve the accuracy of alternatives. The research team specifically concentrate on both uniform and diverse ensemble techniques applied in the context of feature selection. The present investigation of spam detection on Twitter employs user and tweet content characteristics such as age, follower count, and the presence of hashtags and URLs as streamlined features. Four categorization algorithms Support Vector Machines, Logistic Regression, Decision Tree, and Random Forest—were implemented, with Random Forest yielding the highest detection accuracy. The results indicate that using ensemble feature selection improves the effectiveness of machine learning models in identifying spam.