Comparative Study of Machine Learning and Text Vectorization Techniques for Spam Detection

Sai Teja Mantha · 2025

Spam detection remains a critical challenge in natural language processing (NLP) and cybersecurity, with over 50% of global email traffic consisting of unwanted messages. This comprehensive study presents an extensive comparative analysis of machine learning algorithms and text vectorization techniques for spam classification, evaluating seven distinct machine learning models across four feature engineering approaches using multiple large-scale datasets comprising over 15,000 messages. Our experimental results demonstrate that XGBoost achieves the highest overall performance with 94.4% accuracy and 95.4% precision, while ensemble methods consistently outperform traditional approaches by 5-7%. The research reveals that text vectorization techniques show minimal performance variance (less than 0.3% accuracy difference), with Bag of Words (BoW) achieving slightly superior results at 87.9% accuracy. These findings highlight the critical importance of algorithmic sophistication over feature complexity for spam detection systems, providing evidence-based guidance for practical deployment in cybersecurity applications. The study contributes novel insights into ensemble method superiority and establishes comprehensive benchmarks for spam detection research.

Read the paper · More papers on PaperTik