A hybrid ensemble method for spam tweet detection using imbalanced datasets

Tianyu Wang, Yegin Genc, Li‐Chiou Chen · Information Security Journal A Global Perspective · 2025

A series of incidents showed that many security threats caused by Twitter spam tweets can reach far beyond the social media platform and impact the real world. Many studies have applied machine learning techniques to classify spam tweets to alleviate such threats. However, Twitter spam detection faces significant challenges due to class imbalance and the evolving nature of spam techniques. This paper proposes a heterogeneous ensemble approach combining dual LSTM networks with a meta-classifier to address imbalanced Twitter spam detection. Our architecture integrates a similarity-based LSTM processing user behavioral features with a word embedding LSTM analyzing semantic textual patterns, followed by an XGBoost meta-classifier trained on disagreed instances. Experiments on two benchmark datasets (1KS10KN and HSPAM) demonstrate that our model outperforms existing baseline models, achieving F1 scores of 0.952 and 0.945, respectively. The precision-focused approach achieves effective performance suitable for social media content moderation requirements while maintaining practical deployment feasibility. The heterogeneous ensemble framework shows strong potential for application to other imbalanced text classification domains beyond spam detection, providing a robust foundation for cybersecurity and content moderation systems.

Read the paper · More papers on PaperTik