Handling and Research on Imbalanced Data Using VAE-SAGAN ∗
Chenlong Ye, Shaofei Wu, Jun Liu, Zhuoya Hu · 2024
Fraud detection datasets typically contain a large number of normal transaction samples and very few fraudulent samples, resulting in a severely imbalanced dataset. This imbalance affects the model's classification ability, making it prone to misjudgments, which can lead to significant financial losses. Therefore, addressing data imbalances is crucial for improving the accuracy and reliability of fraud detection models. In this study, we propose a fusion model based on Variational Autoencoder and Self-Attention Generative Adversarial Network (VAE-SAGAN) to tackle the data imbalance problem. First, we preprocess the majority of samples using the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) method to filter out overlapping and noisy samples and remove them. Then, we input the minority samples into the VAE-SAGAN model for training, enabling the model to generate more realistic minority samples. In this study, we compare the proposed sampling method with existing sampling algorithms across different datasets and classifiers. Experimental results show that the proposed method outperforms existing sampling algorithms in terms of F1, G-mean, AUC, and Accuracy, and exhibits relatively stable generalization ability across multiple datasets. This research not only enhances the performance of fraud detection models but also provides more reliable technical support for risk management in practical applications.