Classification of imbalanced datasets using partition method and support vector machine
Vinod Kumar Awasare, Surendra Gupta · 2017 Second International Conference on Electrical, Computer and Communication Technologies (ICECCT) · 2017
Classification is a method used to predict target class for each case in datasets. Classification performance depends on the nature of the sample input datasets during the training of classifier. When data samples of one class are more than the data samples of other class, then it is called as imbalanced datasets. In such case, algorithms always favor classifying samples into the overrepresented (majority) class. Existing classification model are inefficient in recognizing samples of the underrepresented (minority) class, which is frequently the class of interest (the minority class). In this paper, Support vector machine with clustering is proposed for learning from imbalanced data in which multiple sub-datasets are built from majority class samples by cluster partitioning method such that each of these sub-dataset has an approximately equal number of samples as a minority class datasets. Then a number of SVM classifiers is developed so that each classifier is trained with a different majority subdatasets and the same minority datasets. Advantage of proposed approach is that there is no loss of valuable information and no overfitting problem because proposed approach is not removing or adding any samples to datasets. Proposed approach is evaluated on many datasets and its performance is compared with existing SVM. Based on an overall result, we can conclude that the proposed SVM with clustering method is very efficient method for class imbalance problem. And result shows that the proposed approach has improved the minority class performance for imbalance problem compared to existing SVM techniques.