A Discriminative Classifier Learning Approach to Image Modeling and Spam Image Identification
Byungki Byun, Chin‐Hui Lee, Steve Webb, Calton Pu · 2007
We propose a discriminative classifier learning approach to image modeling for spam image identification. We analyze a large number of images extracted from the SpamArchive spam corpora and identify four key spam image properties: color moment, color heterogeneity, conspicuousness, and self-similarity. These properties emerge from a large variety of spam images and are more robust than simply using visual content to model images. We apply multi-class characterization to model images sent with emails. A maximal figure-of-merit (MFoM) learning algorithm is then proposed to design classifiers for spam image identification. Experimental results on about 240 spam and legitimate images show that multi-class characterization is more suitable than singleclass characterization for spam image identification. Our proposed framework classifies 81.5 % of spam images correctly and misclassifies only 5.6 % of legitimate images. We also demonstrate the generalization capabilites of our proposed framework on the TREC 2005 email corpus. Multi-class characterization again outperforms single-class characterization for the TREC 2005 email corpus. Our results show that the technique operates robustly, even when the images in the testing set are very different from the training images. 1.