Classification Using Random Forest on Imbalanced Credit Card Transaction Data
Hafija Aktar, Md Abdul Masud, Nusrat Jahan Aunto, Syed Nazmus Sakib · 2021
The aim of the paper is to perform an optimum solution of imbalance classification problem on a real-life scenario like a fraudulent transaction data. The credit card fraudulent datasets are significantly skewed because the number of genuine financial transactions is higher than the number of fraudulent transactions. The term imbalanced dataset refers to a skewed distribution of data objects in which the instances in classes are not evenly distributed. As the distribution of known classes is biased or skewed, instances of one of the two classes are higher than the other. The main objective of this research is to improve testing accuracy in imbalance classification problem. For this purpose, a combination of Random Forest (RF) classifier and repeated stratified k-fold, grid search cross-validation, Synthetic Minority Oversampling Technique (SMOTE), and Random Under-Sampling (RUS) are applied to perform classification. From the experimental results, it is reported that the RF with grid search cross-validation provides the maximum performance in classification accuracy on a highly imbalanced credit card transaction data.