A Selective Learning Model for Spam Filtering.

Didier Colin · 2010

We investigate how one could use optimization methods to attune spam filters to the specific issues and constraints found in the spam filtering area. Among those issues is the need to modelize and manage spammers strategies to delude the filters. To adress this issue, we propose a selective learning scheme designed to maximize learning efficiency. In an offline context, this model uses a simple metaheuristic approach to select a subpart of training data such that the filter induced on that part maximizes its accuracy over the evaluation set. In an online context, we show how one filter can discard incoming messages in order to prevent its knowledge base to be biased by messages which are not good representatives of their class and thus, may lead to a decrease in accuracy. We show that this approach synergizes well with existing classification models while increasing significantly their efficiency over time. More importantly, we show that this model make existing filters less vulnerable to spammers ’ attempts to delude their classification model.

Read the paper · More papers on PaperTik