Evaluation And Comparison of Machine Learning Models for Ham and Spam Email Classification

Sravya Gaddamanugu · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2025

ABSTRACT: Email is one of the most widely used ways of digital communication nowadays, but with that, junk emails have also become more prevalent. These spam messages are not only annoying they can also be harmful, causing security issues such as phishing or data theft. The goal of this research focuses on detecting and filtering spam emails by using machine learning techniques and algorithms. A dataset of 15,267 emails containing spam and ham. Prior to the model training, the dataset was pre-processed by cleaning the textual words converting it into numerical feature vectors using the TF-IDF technique, enabling the algorithms to effectively interpret and analyse the given data. Then applied five well known machine learning algorithms: logistic, svm, naive bayes, decision tree, and random forest. These models were developed and evaluated using tools such as python, the scikit-learn library and Jupiter notebook. To evaluate the effectiveness of each model, performance metrics such as accuracy, precision, recall, F1-score, ROC curve and confusion matrix were employed. Among the models tested, SVM achieved the highest accuracy of closely followed by Random Forest. The results obtained indicates that machine learning models, specifically SVM performs very accurate at detecting spam and improving the security of email communication systems. KEYWORDS: Machine learning, Support vector machine, Spam, Accuracy, Ham, Emails

Read the paper · More papers on PaperTik