Research on random forest algorithm based on oversampling and feature selection
Jiahui Wu, Xiaoxia Lin, Xiaodong Yang, Shuaicai Li, Bingshuo Zhang, Liuyang Gao · 2023
Compared with other classification algorithms, random forest algorithm has obvious advantages in terms of classification accuracy, generalization error and training speed, and is widely used in data mining. However, when faced with unbalanced data, the classification performance of random forest algorithm can be greatly limited. To better deal with imbalanced data, an improved random forest algorithm (oversampling &FeatureSelecctionRandomForest,OF_RF) based on oversampling and feature selection is proposed in the paper. The algorithm firstly, from the data level, uses the Borderline-SMOTE algorithm to preprocess the imbalanced data set and introduces the LOF algorithm to improve the Borderline-SMOTE algorithm in order to eliminate outliers and thus improve the processing performance for negative class samples. Secondly, from the algorithm level, the ReliefF algorithm is used to assign different weights to the balanced processed data, while excluding irrelevant and redundant features, so as to perform dimensional simplification. Finally, the classification performance of the algorithm is further improved by using the weighted voting principle. The experimental results show that the improved OF_RF algorithm exhibits higher evaluation metrics in processing unbalanced data compared with the traditional random forest algorithm, proving that the algorithm has significant advantages in classification performance of unbalanced data.