Comparative Analysis of Machine Learning Algorithms through Credit Card Fraud Detection
R. Banerjee, Gabriela Bourla, Steven Chen, Mehal Kashyap, Sonia Purohit · 2018
With the increase of e-commerce and online transactions throughout the twenty-first century, credit card fraud is a serious and growing problem. Such malicious practices can affect millions of people across the world through identity theft and loss of money. Data science has emerged as a means of identifying fraudulent behavior. Contemporary methods rely on applying data mining techniques to skewed datasets with confidential attributes. This paper examines numerous classification models trained on a public dataset to analyze correlation of certain attributes with fraudulence. This paper also proposes better metrics for determining false negatives and measures the effectiveness of random sampling to diminish the imbalance of the dataset. Finally, this paper explains the best algorithms to utilize in datasets with high class imbalances. It was determined that the Support Vector Machine algorithm had the highest performance rate for detecting credit card fraud under realistic conditions.