Using GNUsmail to Compare Data Stream Mining Methods for On-line Email Classification
José M. Carmona-Cejudo, Manuel Baena-García, Rafael Morales Bueno, João Manuel Portela da Gama, Albert Bifet · 2011
Real-time classication of emails is a challenging task because of its online nature, and also because email streams are subject to concept drift. Identifying email spam, where only two dierent labels or classes are dened (spam or not spam), has received great attention in the literature. We are nevertheless interested in a more specic classication where multiple folders exist, which is an additional source of complexity: the class can have a very large number of dierent values. Moreover, neither cross-validation nor other sampling procedures are suitable for evaluation in data stream contexts, which is why other metrics, like the prequential error, have been proposed. In this paper, we present GNUsmail, an open-source extensible framework for email classication, and we focus on its ability to perform online evaluation. GNUsmails architecture supports incremental and online learning, and it can be used to compare dierent data stream mining methods, using state-of-art online evaluation metrics. Besides describing the framework, characterized by two overlapping phases, we show how it can be used to compare dierent algorithms in order to nd the most appropriate one. The GNUsmail source code includes a tool for launching replicable experiments.