Deep Dive on Oversampling and Under Sampling Techniques in Machine Learning

K. Vijay, J. Manikandan, Babu Rajendiran, K. R. Sowmia, Eugene I. Berna · 2023

Data imbalance is a term used in Computer Vision to refer to an unbalanced distribution between classes throughout a dataset. In classification tasks, where the dispersion of classes or labels in a given data is not regular, this problem is most frequently encountered. The resampling approach, which involves adding data to the minority class and removing entries from majority class, is the most feasible approach for solving this problem. In this study, we conducted experiments with two frequently used resampling techniques: oversampling and under sampling, both of which have been widely accepted. In order to investigate both resampling methodologies, we used a publicly available imbalanced dataset from the Kaggle website and a collective of very well ML algorithms with distinct hyperparameters that produced the best outcomes for both resampling methodologies. One of the most important discoveries of this study is the observation that oversampling outperforms under sampling for various classifiers and results in increased scores in many evaluation metrices.

Read the paper · More papers on PaperTik