A Comprehensive Evaluation of Over-Sampling and Under-Sampling Methods for Imbalanced Data in Anomaly Detection

Raghav Juneja, Pooja Pooja, Shivaansh Pande, Maadhav · 2025

Credit card fraud plagues banks and other financial institutions. It costs them billions of dollars every year. Foundational methods struggle to detect difficult patterns in financial data thus there is a need of sophisticated algorithms which is solved by Machine Learning (ML) and Deep Learning (DL) based approaches for fraud detection. These models continuously keep on improving and learning from new fraud scenarios and thus generalizing a lot better and achieving a low false positive rate. In this paper, authors have analyzed fraud using a comprehensive European dataset spanning two days with fraud and non-fraud transactions, with the only problem being non-fraud i.e. legitimate transactions constitute approximately 99.9% of all cases. In this paper, authors initially used an approach without any resampling techniques and since the dataset was severely imbalanced, the accuracy achieved was high although the model didn't generalize better to fraud transactions and did overfit. To address this challenge, authors explored robust resampling techniques under over-sampling and under-sampling along with certain ML models to devise an approach which is easily able to detect a fraud transaction by learning the sophisticated features in the dataset. The models were evaluated on an industry standard metrics such as precision, recall, f1 Score and ROC-AUC curve with particular emphasis on improving false positive rates under hyper parameter tuning. Further, this paper went one step forward to introduce ensemble techniques too. Undersampling, however leads to loss of data which can be detrimental for credit card fraud detection, as it reduces the available information needed. So, in stage 3 authors tried using oversampling techniques. This paper used over 8 over sampling techniques combined with a variety of ML models followed by hyper parameter tuning to select the ones which outperform others.

Read the paper · More papers on PaperTik