Research on the Classification of High Dimensional Imbalanced Data Based on the Optimizational Random Forest Algorithm

Bo Su · 2017

The random forest algorithm is a new classification and prediction model algorithm. So far, there is not much research on the problem of unbalanced data for random forest classification, ditto, no direct and effective method. On the basis of feature selection algorithm based on correlation measure, the integration feature selection method was helpful to increase the selection probability of classification feature of positive class samples, which used the subspace selection algorithm based on stratified sampling, sampling the feature subset generated by the integration feature selection method, respectively, which ensured the importance of the feature and the difference of the generated model. The simulation results show that the proposed method is effective.

Read the paper · More papers on PaperTik