Unbalanced Data Processing and Machine Learning in Credit Card Fraud Detection

Haili Peng, Jing Wang · Research Square · 2022

Abstract Imbalanced data often performs poorly in the model and can prevent the model from capturing a few classes of samples, so it is crucial to process imbalanced data. This paper is a credit card fraud detection based on imbalanced data, comparing different processing methods for imbalanced data and using machine learning to detect credit card fraud, and finally arriving at optimal results. Since credit card fraud data is mostly a dichotomous problem and highly imbalanced, it means that the machine learning model favors the majority of the samples and treats the fraudulent transactions in the credit card fraud data as correct transactions. The treatment of imbalanced data is crucial because of the low percentage of fraudulent data. We used different methods for imbalanced data such as oversampling, undersampling, combined sampling and using class weights to improve the class imbalance and applied these methods to credit card fraud detection and calculated Accuracy, Precision, Recall, F1 score and AUC. Because of the severe imbalance in the data, the model is biased towards majority of the samples, so the accuracy of the model will be high. Because we focus more on the probability that a minority class of the sample is correctly classified, to check the performance of the model, we will use the F1 score, the Area Under the Precision-Recall Curve (AUPRC). and recall as measures instead of accuracy. The results demonstrate that the model achieves the best performance by resampling the credit card fraud data, and finding the optimal weights for different category samples by grid search when setting the category weights leads to a significant improvement in the performance of the logistic regression model, and the random forest outperforms all the machine learning models that are compared.

Read the paper · More papers on PaperTik