Collective classification for spam filtering
Carlos Laorden, Borja Sanz, Igor Santos, Patxi Galán-García, Pablo G. Bringas · Logic Journal of IGPL · 2012
Spam has become a major issue in computer security because it is a channel for threats such as computer viruses, worms and phishing. Many solutions to the spam problem feature machine-learning algorithms that are trained using statistical representations of terms that often appear in spam e-mails. However, these methods require a training step with labelled data. Dealing with situations in which the availability of labelled training instances is limited slows the filtering systems' progress and offers advantages to spammers. Currently, many approaches direct their efforts at Semi-Supervised Learning (SSL). SSL is a halfway method between supervised and unsupervised learning. In addition to using unlabelled data, SSL receives supervision information such as associations between targets and examples. Collective Classification for Text Classification is an interesting method for optimizing the classification of partially labelled data. Here, we propose for the first time the use of Collective Classification algorithms for spam filtering to overcome the amount of unclassified e-mails that are sent every day.