A fusion algorithm model based on KNN-SVM to classify and recognize spam

Zengle Ge · Journal of Physics Conference Series · 2021

Abstract Spam usually has the characteristics of large data dimension and large numbers of samples. At present, the common classification learning algorithms have their own advantages and disadvantages. Aiming at the classification tendency of K nearest neighbor algorithm (KNN) in the face of unbalanced sample data and the long time consuming of support vector machine (SVM) training model, a fusion algorithm of KNN and SVM is proposed. Firstly, the content of spam is extracted, and then the text preprocessing is transformed into sparse matrix. The k nearest neighbor samples of the test sample set are screened out by KNN algorithm, and then the SVM algorithm is used to train these screened samples, establish the separation hyperplane, identify the mail, and judge whether it is spam or not. The results show that the KNN-SVM fusion algorithm can effectively reduce the impact of sample imbalance while maintaining the high classification accuracy of SVM algorithm, and its classification efficiency is slightly lower than that of KNN, algorithm but much higher than that of SVM algorithm.

Read the paper · More papers on PaperTik