Ham-Spam Filtering Using Different PCA Scenarios
Issam J. Dagher, Rima Antoun · 2016
The objective of this paper is to discuss different scenarios for Principal Component Analysis classifier implemented for email filtering process (Ham vs. spam emails). The study highlights on the variation of the accuracy of these classifiers with respect to the variation in feature preprocessing. Four scenarios were considered: Scenario 1: Ham and Spam classes are represented with different features. Scenario 2: Ham and Spam classes are represented with same features. Scenario 3: Ham and Spam classes are represented with common terms. Scenario 4: Ham and Spam classes are represented with common Features and Characteristic terms. Different experiments were done using a public corpus extracted from the University of California-Irvine Machine Learning Repository. Different training and test sets were used. A comparison with Support Vector Machine and Bayes detector was done to prove its superior behavior.