Impacts of data preprocessing and TF-IDF based feature extraction on chinese spam recognition in machine learning methods
Jun Gao, Tian Qiu, Guihai Chen · IET conference proceedings. · 2023
Spam has been threatening people's life and production safety, in order to reduce its harm, it is necessary to study the topic of spam filtering. In spam identification, data preprocessing and feature extraction are two important aspects. In this paper, the impact of combing TF-IDF (Term Frequency-Inverse Document Frequency) based feature extraction and different preprocessing methods on Chinese spam recognition is investigated. We analyzed the effectiveness of mainstream machine learning methods on Chinese spam identification with or without data preprocessing. We discussed in detail the effect of different sizes of features on the classification results. The SVM (Support Vector Machine) model presented in this paper achieves very high accuracy (99.0%) and F1 (99.26%) value in spam filtering. It shows that data preprocessing improves the performance of machine learning models, and the degree of improvement is related to the selected model. Experiments have shown that to achieve the same accuracy, the random forest model requires fewer features than the Bayesian model.