Comparison of Undersampling Techniques in Credit Card Fraud Modeling Cases
Vio Greacely Simatupang, Tora Fahrudin, Renny Sukawati · 2025
In this era, online fraud often occurs, such as credit card fraud. Therefore, various models and methods are used to detect fraudulent activities. This research was conducted to determine the best models and methods for detecting credit card fraud. Fraud detection is carried out by identifying patterns, anomalies, and suspicious behaviors from the dataset used. The dataset currently used is imbalanced, so the undersampling method is employed to balance the classes of non-fraudulent and fraudulent data, after which the machine learning model will be trained. The undersampling methods include Random Under Sampler, Cluster Centroids, NearMiss, Tomek Links, and Edited Nearest Neighbors. The classification models used include Decision Tree, K-Nearest Neighbors, Support Vector Machine, Artificial Neural Network, and Naive Bayes. This research uses the method of addition and division. The metric results obtained from the experiments will be summed and then divided to get the average. The results of this study show the top 3 evaluations: Decision Tree at 0.832 or 83.2%, SVM at 0.66 or 66%, and Naive Bayes at 0.598 or 59.8%. The research indicates that the best model is the Decision Tree with a precision of 0.73, a recall of 0.79, an F1-score of 0.75, and an AUC-ROC of 0.89. From the best model, the best method obtained is Edited Nearest Neighbors at 0.6424 or 64.24%.