Benchmarking Machine Learning Techniques for Credit Card Fraud Detection
Dhwanir Shah, Lokesh Kumar Sharma · Indian Journal of Science and Technology · 2025
Objective: An analysis aimed at measuring the success of five machine learning algorithms—Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, and XGBoost—in recognizing credit card fraud, based on a Kaggle dataset offered by Mr. Anurag Verma for credit card transactions. Methods: The dataset was segmented into an 80% training set and a 20% testing set. To counteract class imbalance, the Synthetic Minority Over-sampling Technique was applied. The models' performance was assessed using Accuracy, Precision, Recall, F1-score, ROC-AUC, and PRUAC, with validation executed via K-fold cross-validation for K = 3, 5,10. Findings: XGBoost exhibits a commendable equilibrium in its performance metrics, achieving a precision of 84.44%, a recall of 73.57%, and an F1-score of 78.63%. Notably, it excels with the highest PRAUC of 88.65%, demonstrating exceptional performance in contexts with potentially unbalanced classes. Other algorithms may present marginally superior values in specific metrics, the amalgamation of XGBoost's competitive precision, recall, and F1-score, alongside the leading PRAUC, indicates a strong and dependable model. For example, even if an alternative algorithm boasts a higher recall, XGBoost sustains a more favourable balance between precision and recall, as evidenced by its F1-score and PRAUC. Its impressive accuracy of 97.12% and ROCAUC of 98.77% further validate its efficacy. Novelty: This research outlines several essential methodological considerations. We implement k-fold cross-validation, varying k at 3, 5, and 10, to assess the robustness and generalization performance of the model. GridSearchCV is utilized for comprehensive hyperparameter tuning to optimize the model's performance. A data split of 80:20 for training and testing is adopted. We incorporate PRAUC as one of the evaluation metrics, with a particular emphasis on the performance of the positive class. Additionally, we apply both Interquartile Range and Z-score methods for effective outlier detection, and we employ a Random Forest method for feature selection. Keywords: OneHotEncoding, PRAUC, SMOTE, IQR, Z-Score, Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, XGBoost, K-Fold cross validation