SHRED: An Ensemble-Based Machine Learning Model to Sift Email Messages for Real-Time Spam Detection

Shahid Alam, Amina Jameel, Zahida Parveen, Ehab Tawfeek Alnfrawy · IEEE Access · 2025

The rapid expansion of email usage has been paralleled by a significant rise in spam, which not only clutters inboxes but also spreads malware and introduces serious security threats. While machine learning (ML) has proven to be a powerful approach for spam detection, single-classifier models often face issues such as poor generalization, overfitting, and elevated false positive rates. Key challenges in email spam detection include: (1) Improving accuracy while keeping false positives low to maintain usability; (2) Responding to threats in real-time; (3) Handling the ever-changing characteristics of email content; and (4) Automating model adaptation to these changes. To address these challenges, this study presents SHRED (Sifting Ham for Real-time Email Spam Detection), an ensemble-based ML model specifically built for real-time spam filtering. SHRED combines Naive Bayes, Decision Trees, Adaptive Boosting, Random Forests, and Artificial Neural Networks using a hybrid approach of voting and stacking to enhance detection performance. It employs preprocessing methods such as text normalization and TF-IDF-driven feature selection to produce resilient input representations that automatically adapt to evolving email patterns. When tested on a diverse dataset of 83,448 emails, SHRED achieved a detection rate of 98.4%, a false positive rate of 1.6%, and an overall accuracy of 98.45%, with an average processing time of just 0.0522 seconds per email. This includes all stages: preprocessing, feature extraction and selection, training, and detection, underscoring the system’s capability to accurately detect spam in real-time.

Read the paper · More papers on PaperTik