Analysis of Machine Learning Models for Spam Email Detection and Real-Time Integration

Ezinne Ruth Ejirika, Temidayo Oluwatosin Omotehinwa · 2024

Spam fraud has surged recently, posing a global challenge. It inflicts significant financial harm through email scams, necessitating urgent action. To address this, we developed a machine learning model for spam email detection. We utilized a subset of the Enron1 dataset comprising 7,000 spam and legitimate emails. Employing R, we trained Classification and Regression Tree (CART), Support Vector Machine (SVM), Naïve Bayes (NB), and Random Forest (RF) algorithms on a dataset of 2,614 legitimate emails and 2,309 spam emails. We evaluated model performance based on various metrics including sensitivity, accuracy, specificity, precision, F1-score, and area under the ROC curve. The RF model emerged as the top performer, achieving high scores across all metrics: sensitivity (0.9735), accuracy (0.9839), specificity (0.9641), precision (0.9615), F1-score (0.9726), and ROC value (0.9740). Integrating this superior RF model into a real-time mail management system developed using HTML, CSS, PHP & MySQL enabled accurate classification of emails, displaying the probability of each mail being spam or legitimate. While the RF model shows promise, fine-tuning its hyperparameters, such as ntree and mtry, is recommended to enhance prediction accuracy further.

Read the paper · More papers on PaperTik