A Comparison Study of Classification Methods and Effects of Sampling on Unbalanced Data

Atharva Deshpande, Chinmay Kamath, Mihir Joglekar · 2019 International Conference on Smart Systems and Inventive Technology (ICSSIT) · 2019

Most of the real-world classification problems exhibit an amount of imbalance of class. With such data sets, an algorithm doesn't get the required information about the minority class to make a correct prediction. Hence, imbalanced classification becomes very important. This paper explores different methods for imbalanced classification which is a supervised learning problem in which one class is larger than the other. This problem occurs often in binary classification problems than multi-level classification problems. An imbalanced classification problem is in which the dependent variable does not have a balanced proportion of classes. We use one such highly imbalanced dataset (made up of credit card transactions carried out in 2 days during September 2013 by customers of a European bank) to study and compare existing classification models by trying to predict fraudulent transactions from the legit ones. We also look at how different sampling methods like undersampling and oversampling will impact the prediction rate of these classifiers and thus try to create an optimal sampling strategy for this dataset. we try to evaluate performance based on new performance metrics like Precision (the measure of correctness achieved in positive prediction), and Recall (the ratio of observations predicted correctly). This paper explores a comparison study between common classification models like Logistic Regression, KNN SVM and Decision Tree Learning and the effects of various sampling rates on unbalanced data.

Read the paper · More papers on PaperTik