Improving Credit Card Fraud Detection Through Stacking Ensemble Models and SMOTE-ENN for Imbalanced Datasets
Er. Jyoti, Chirag Jindal, Naman Deep Singh, Akhil Saini · 2024
Credit card fraud detection continues to be an open problem especially since datasets are usually imbalanced in the context of the financial industry. In this work, the stacking ensemble model along with SMOTE-ENN, which is the Synthetic Minority Over-sampling Technique with Edited Nearest Neigh-bors, is proposed to maintain the proportion between fraudulent and non-fraudulent transactions. We used a relatively big 1 as the dataset to achieve good results during the identification of significant connections among different variables. Transactions focused on, 85 million records including features like; transaction amount, geo coordinates, city population and merchant locations. The data cleaning included treatment of missing values, mean imputation, normalization of features to unit variance using StandardScaler and finally feature selection by IPCA. The stacking ensemble consisted of Random Forest, XGBoost and Gradient Boosting as base models; while Logistic Regression was used as the meta-classifier. Thanks to the implementation of a number of performance indicators, the model produced an overall accuracy of 0.97 while their ROC AUC score was 0. 9862. However, the above model was able to identify only 7496 and out of 7506 fraudulent transactions while creating 33739 false positives and therefore had a precison of 0.18 for the fraud class. This work exposes the conflict of interest between accuracy and coverage for imbalanced datasets and outlines possibilities for future enhancement.