Machine Learning Approaches for Predicting Bank Customer Subscription: A Comparative Analysis
Zibo Zhao · ITM Web of Conferences · 2025
As people enter the big data era, the traditional banking industry faces huge competitive pressure from new Internet financial products, which requires the traditional banking industry to use data mining and machine learning methods to optimize marketing strategies. Based on the bank marketing data set, this paper explores the effectiveness of several machine learning methods in predicting potential customers, including random forest, k-nearest neighbor (KNN) algorithm, and logistic regression, and conducts comparative analysis before and after data oversampling. Experiments show that the synthetic minority oversampling technique can effectively strengthen the model's capacity to recognize minority samples, but it may cause overfitting. Among them, the random forest has the best overall performance; logistic regression is limited by its linear assumption and performs slightly worse; the KNN algorithm is sensitive to noise and unbalanced data and has poor results. Future research can explore combining different sampling methods or using model integration to improve performance. This study provides important reference significance for the banking industry in accurately positioning potential customers.