Automatic Classification and Analysis of Spam Based on Machine Learning
Jiacheng Xu · 2023
With the proliferation of Internet use and online communication, the issue of spam, characterized by the inundation of unsolicited or unwanted messages in large numbers, has become an enduring challenge. Spam classification, the automated task of identifying and filtering out spam, has been a subject of extensive research. This paper evaluates the performance of five machine learning models, namely Naive Bayes, Support Vector Machines (SVM), Logistic Regression, Random Forest, and LightGBM, through comprehensive training using confusion matrices and accuracy ratings. The evaluation process is carried out following three main steps. First, the spam dataset is meticulously analyzed and subjected to pre-processing, incorporating feature extraction techniques to adapt its data format. Second, the five machine learning algorithms are applied to construct classification models for data training. Lastly, the strengths, weaknesses, and comparative performance of these algorithms are rigorously assessed. The experimental findings highlight the effectiveness of Naive Bayes and SVM in spam classification. This methodology is instrumental in the development of efficient spam classifiers and enhancing email filtering processes. Moreover, the paper discusses potential future directions in the field aimed at advancing and refining spam detection techniques.