Towards a Smarter Spam Filter: Leveraging Machine Learning for Accurate Email Classification
Vinod Jain · 2025
With the ever-increasing reliance on electronic communication, the proliferation of unsolicited spam emails has emerged as a significant challenge for both individuals and organizations. These spam messages not only clog inboxes but can also pose serious threats in the form of phishing, malware distribution, and identity theft. This research aims to develop a robust spam email classification system using machine learning techniques to automatically distinguish between spam and legitimate messages. The classifiers implemented and compared in this study include the Multilayer Perceptron (MLP), Multinomial Naive Bayes (MNB), and Bernoulli Naive Bayes (BNB) algorithms. The dataset used comprises 5572 pre-labeled email messages, enabling supervised learning. The models were trained and evaluated using standard performance metrics such as accuracy, precision, recall, and F1-score. The experimental results demonstrate that all three classifiers achieve high accuracy, with the MLP classifier slightly outperforming others in overall performance. This work presents an effective email spam detection system by evaluating MLP, Multinomial Naive Bayes, and Bernoulli Naive Bayes classifiers, where the MLP achieved the highest accuracy (0.99), while Naive Bayes models offered faster, recall- and precision-optimized alternatives for real-time applications.