A spam classification method based on NB and SVM

Liyun Li · 2023

The dataset used in this project is derived from the SMS spam classification dataset in the UCI Dataset Repository, and it is necessary to understand what the dataset looks like before pre-processing the data. The first step is text clean-up, and the second step is text feature extraction. This paper investigates and compares the performance of classifiers combining different dimensionality reduction methods on spam datasets to provide a reference for related classification studies. The project then uses the scikit-learn machine learning library to train the classifier, dividing the dataset into 75% training sets and 25% test sets, and introducing classifiers such as NB, IR, SVM for training. After classifier training is complete, test the result of the model on the test set. Use trained classification models to predict the category of a message (regular mail or spam) The result shows that the best performer among the various classifiers is the SVM.

Read the paper · More papers on PaperTik