Using Genetic Algorithm to Improve Classification Accuracy on Imbalanced Data

Jair Cervantes, Xiaoou Li, Wen Yu · 2013

Many real data sets are imbalanced, which contain a large number of certain type objects and a very small number of opposite type objects. Normal classification methods, such as support vector machine (SVM), do not work well for these skewed data sets. In this paper we propose a genetic algorithm (GA) based classification method. We first use SVM to generate a draft hyper plane and support vectors. Then GA is applied to find new data points in the sensible region or classification margin. Finally, SVM is used again to find the best hyper plane from the generated data points. Compared with the other popular classification algorithms, the proposed method has better classification accuracy for several skewed data sets.

Read the paper · More papers on PaperTik