Neighborhood Triangular Synthetic Minority Over-sampling Technique for Imbalanced Prediction on Small Samples of Chinese Tourism and Hospitality Firms
Yu Hui Xu, Hui Li, Lu Le, Xiao Yun Tian · 2014
In order to solve the problem of unsatisfactory results of imbalanced risk prediction on minority class samples, we suggested to adjust the up-sampling approach to be the neighborhood triangular synthetic minority over-sampling technique (NT-SMOTE). The new approach that we add the nearest neighbor idea and the triangular area sampling idea to the SMOTE performed better in dealing with samples of minority class by turning imbalanced problems into balanced ones. Thus, performance of single classifiers in predicting risk on imbalanced and small datasets was improved. By using the related knowledge of data excavation principles, the data of listed companies of the Chinese tourism and hospitality industry were processed. Missing samples and missing financial indicators were eliminated. Significant indicators of financial data were filtered out with significance test. Then, NT-SMOTE was used to over-sample minority samples. Further, we used a variety of popular single classifiers of financial risk prediction, including: MDA, DT, LSVM, Logit, and Probit, for risk prediction. These single classifiers improved with NT-SMOTE can reasonably and effectively solve the problem of imbalanced and small sample oriented firm risk prediction.