Evaluation of Supervised Learning Models for Automatic Spam Email Detection

Tsehay Admassu Assegie · Research Square · 2023

Abstract This paper compares the performance of different supervised learning algorithms for email spam detection. The comparison considered performance measures such as the area under the curve (AUC), F-score, precision, and confusion matrix. The paper evaluated the performance of eight supervised learning algorithms for email spam detection. The first stage collected the dataset of emails from the Kaggle repository. The second stage involves pre-processing, duplicate removal, and the dataset features scaling. After the pre-processing stage, the study employed a synthetic minority technique for balancing samples representing the spam and no spam emails in the dataset. In the final stage, the supervised learning algorithms are trained on the pre-processed dataset and then the test result is analyzed. The comparison shows the random forest (RF) model performing at higher accuracy than the other models. The RG model achieved 96.6% accuracy in email spam detection. Thus, the result demonstrates that different models tend to perform differently in email spam detection.

Read the paper · More papers on PaperTik