Personalized Spam Filtering with Natural Language Attributes
Rushdi Shams, Robert E. Mercer · 2013
Email spam is one of the biggest threats to today's Internet. To deal with this threat, many anti-spam filters have been developed. One big challenge for these filters is to predict the labels of emails in a personalized mailbox. In this paper, we report the performance of an anti-spam filter named Sentinel. In addition to some commonplace attributes, Sentinel uses attributes related to natural language stylometry. The filter has been tested with six benchmark datasets in the Enron-Spam collection. Classifiers generated by well-known meta-learning algorithms like AdaBoostM1 and Bagging perform equally the best, while a Random Forest (RF) generated classifier performs almost as well. The performance of classifiers using Support Vector Machine (SVM) and Naive Bayes (NB) are not satisfactory. Comparisons show that the performance of Sentinel surpasses that of a number of state-of-the-art personalized filters proposed in previous studies.