Research on Classification Method of High-Dimensional Class-Imbalanced Data Sets Based on SVM
Chunkai Zhang, Jianwei Guo, Junru Lu · 2017
In recent years, the problem of classification for high dimensional and class-imbalanced data is found in many fields like bioinformatics and so on. High dimensional problem result in bad classification results because of some combinations of features have adverse effect on classification. Class-imbalanced problem means the number of samples of one class is more than another class, which would make the classifier concerns the majority class more but the minority less. The two problems are both exist in high dimensional and class-imbalanced data sets. Many researchers make researches on high dimensional problem and class-imbalanced problem separately and come up with a series of algorithms. They ignored the new problem arising from the mutual influence of class-imbalanced problem and high dimensional problem. This article introduces the two problems and analysis the new problem arising from the influence of the two problems firstly. And then this article introduces SVM, analysis its advantages on dealing high dimensional problem and class-imbalanced problem. Next, this article improves SVM-RFE by considering the class-imbalanced problem in the process of feature selection and improve SMOTE so that the procedure of over-sampling could work in the Hilbert space and the over-sampling rates are set adaptably meanwhile. Finally, a classification algorithm aimed at high dimensional and class-imbalanced data sets is come up in this article which named BRFE-PBKS-SVM: Border-Resampling Feature Elimination and PSO Border-Kernel-SMOTE SVM. And a series of experiments were made to prove the effectiveness of this algorithm by using different evaluation indexes.