Classification Models and Variable Selection for SME Credit Risk Assessment in Balance the Data Set
Sornthun Intharasompong, Anuchit Jitpattanakul, Sakorn Mekruksavanich, Piyada Wongwiwat, Wikanda Phaphan · 2025
Small and medium-sized enterprises (SMEs) are frequently considered high credit risk, making it difficult to obtain loans from banks. A viable substitute for assessing borrower creditworthiness more impartially is using machine learning-based credit risk assessment (CRA) algorithms. This study evaluated the performance of five machine learning techniques: decision tree, support vector machine, gradient boosting, K-nearest neighbor, and Naïve Bayes. We evaluated each model with K-fold cross-validation and utilized the Stepwise selection method for the variable selection. To address the underlying data imbalance, we balanced the dataset with the SMOTE technique, which showed that normal loans were more widespread than Non-Performing Loans (NPLs). After balancing the data, the Gradient Boosting model displayed the best accuracy, F-measure, and robustness without variable selection, making it the most effective model for predicting SME loans.