Email Spam Detection with Machine Learning and Vectorization

R Saravanan, Vaisshale Rathinasamy · 2024

Addressing the current pressing issue of managing unwanted and potentially hazardous email communications, this study focuses on the Ling spam dataset to effectively categorize spam emails. To address the dataset’s imbalance, the study employs ADASYN, an oversampling technique, to rectify the biased distribution. Following this, two vectorization models, FastText and GloVe, were employed, and three machine learning algorithms, Random Forest, Navies Bayes and Logistic Regression, were utilized for spam/ham categorization. Among the vectorization models, GloVe stands out as the optimal choice. Notably, when combined with Naive Bayes, GloVe demonstrates the best performance, achieving an impressive accuracy rate of $\mathbf{9 8. 7 \%}$ on the Ling spam dataset. This surpasses existing research works, including a prior study by Samira Douzi [1], which employed an integrated approach to spam filtering combining various techniques, achieving a $98.27 \%$ accuracy rate.

Read the paper · More papers on PaperTik