Hybrid GA-SVM for Efficient Feature Selection in E-mail Classification
Fagbola Temitayo, Olabiyisi Stephen Adigun Abimbola · Computer Engineering and Intelligent Systems · 2012
Feature selection is a problem of global combinator ial optimization in machine learning in which subse ts of relevant features are selected to realize robust le arning models. The inclusion of irrelevant and redu ndant features in the dataset can result in poor predicti ons and high computational overhead. Thus, selectin g relevant feature subsets can help reduce the comput ational cost of feature measurement, speed up learn ing process and improve model interpretability. SVM classifier has proven inefficient in its inability to produce accurate classification results in the face of larg e e-mail dataset while it also consumes a lot of computational resources. In this study, a Genetic A lgorithm-Support Vector Machine (GA-SVM) feature selection technique is developed to optimize the SV M classification parameters, the prediction accurac y and computation time. Spam assassin dataset was used to validate the performance of the proposed syste m. The hybrid GA-SVM showed remarkable improvements over SVM in terms of classification accuracy and computation time.