Empirical Analysis of Sampling Methods on Imbalanced Data

N Disha, Neha B Chadaga, Hemang Singh, N Kayarvizhy · 2022

With imbalanced data being one of the significant drawbacks encountered while evaluating the effectiveness of machine learning models, where the distribution of instances within a dataset is skewed, we aim to provide data-level solutions to manipulate the balance of data before feeding them onto classifiers. The decline in performance in such datasets occurs when the accuracy in minority classes is poor due to very few instances available in them to train the model. In this paper, we assess various existing sampling techniques in order to handle skewed data to enhance the classification performance of the model which includes traditional Oversampling and Undersampling and five variants of Smote: Smote, Borderline Smote, SVM Smote, Adasyn, and SmoteTomek. We have experimented with these methods on a variety of imbalanced binary classification datasets ranging from highly imbalanced to moderately imbalanced and evaluated them based on performance metrics such as recall, precision, and F-measure.

Read the paper · More papers on PaperTik