The Effects of Sampling Technique and Sample Size on Classification Performance of Ensemble Tree Classifiers in Detecting Credit Card Fraud
Mcdonald Chogugudza, Sikhumbuzo Ngwenya, Khulumani Sibanda · 2024
Despite the implementation of fraud detection and prevention systems, credit card fraud continues to pose a significant threat due to the constantly evolving tactics employed by fraudsters. Thus, there is a pressing need for ongoing research to mitigate losses stemming from fraudulent activities. Despite the myriad classifiers utilized in this field, little attention has been given to assessing the efficacy of ensemble tree classifiers in combating fraud. To address this gap, this study employs an experimental methodology to evaluate the effectiveness of ensemble tree classifiers in identifying credit card fraud, along with examining how various parameters in modeling fraud detection systems impact their classification performance. Specifically, the study investigates the influence of sample size and sampling technique. Results indicate that larger sample sizes contribute to superior classifier performance, attributed to models having more data instances to learn from, thereby encompassing a wider array of fraudulent transaction patterns. Additionally, the choice of sampling technique emerges as a critical factor affecting classification performance; undersampling demonstrates a notable reduction in performance, whereas oversampling leads to improved outcomes.