Experimental study of spam classifier based on naive Bayesian model

Teng Lv, Ping Yan · 2023

These days, we are bombarded with endless spam e-mails. A survey showed that spam accounts for over 70% of all emails. The harm of spam to our daily life includes: spam usually has fraud and unhealthy contents, which requires a vast waste of network bandwidth to transfer spam, and vast waste of space to store spam. As spam is usually embedded in normal e-mails, it is difficult to identify them. In this paper, main technologies to identify and block spam are analyzed including information filtering, blacklist and white list, and intention analysis. Then two experiments of different settings are conducted and analyzed to show how different settings affects the accuracy, precision, recall, and f1_score of the model: the first experiment shows how different thresholds are set to determine whether an e-mail is a spam or not effects the accuracy, precision, recall, and f1_score of the model, the second experiment shows the effects on accuracy, precision, recall, and f1_score of the model when training set is bigger than test set.

Read the paper · More papers on PaperTik