To improve classification of imbalanced datasets

Pratyusha Shukla, Kiran Bhowmick · 2017

The task of accurately predicting the target class for each case in the data is called classification of data in data mining. Classification of balanced data set is fairly simple and easy to perform but it becomes difficult when the data is not balanced. Class Imbalance problem is the problem in machine learning where the total number of a class of data (positive) is far less than the total number of another class of data (negative). In this paper, we have used K-Means algorithm to balance the imbalanced dataset and then use SVM to classify the balanced dataset. We have compared the accuracy, precision, recall and time taken in classifying balanced as well as imbalanced datasets and results show that K-means helps in balancing the data and hence the accuracy and time taken to classify balanced dataset is much better than simply classifying the imbalanced dataset.

Read the paper · More papers on PaperTik