Comprehensive Email Spam Detection Using Large Language Models: Exploring the Limits of LLM-Based and Traditional Methods

Aya Salama Abdelhady, Mohammad Shaker · 2025

Spam email remains a massive problem despite better spam filtering, rule-based and machine learning systems find it hard to keep up with spammers evolving techniques. While Large Language Models (LLMs) have been found strong in text classification, the application of a single LLM has several drawbacks, including inconsistency in classification, vulnerability to manipulation by other LLMs, and inability to handle diverse patterns of spam. To solve these problems, this paper investigated the efficiency of ensemble LLMs for spam detection. By using more than one model, ensemble methods can improve robustness, increase classification accuracy, and eliminate biases that come with a single model. The combination of BERT, RoBERTa and DistilBERT delivered the highest accuracy of 98. 46%, and an F1 score of 98. 47%, these remarkable results were achieved with a training time of only 32.82 seconds, highlighting the efficiency of the proposed method. This efficiency not only ensured faster deployment but also made the ensemble-based approach a practical solution for real-time spam classification in production environments, making it the most effective configuration for balanced spam detection.

Read the paper · More papers on PaperTik