Synthetic Data Generation and Handling Data Imbalance for Mobile Financial Transactions

Abhilash Mohapatra, Abhijeet Kumar, Bimlesh Kumar, Harshit Agarwal, Rojalina Priyadarshini · 2024

It is always challenging to get the financial transaction data due to the data privacy issue. This research paper presents an approach for detecting financial fraud using machine learning techniques by augmenting the data. The proposed method leverages the power of synthetic data generation techniques to create large volumes of realistic data that can be used to train and evaluate machine learning models. Specifically, here Generative Adversarial Networks (GANs) are used to generate synthetic data that closely resembles real-world data while preserving its statistical properties. The generated data are highly imbalanced. So measures are taken to handle the data imbalance. Moreover, the noises present in the data are also removed by utilizing edited nearest neighbourhood techniques. Then these data are used to train and test several machine learning models, including Random Forest, Support Vector Machines (SVM), and logistic regression. Our results show that using synthetic data can significantly improve the performance of machine learning models for fraud detection, achieving high levels of accuracy, precision, recall, and F1-score. Moreover, our approach is shown to be robust to imbalanced datasets, which is a common challenge in fraud detection.

Read the paper · More papers on PaperTik