Random Forest-Based Analysis of Variability in Feature Impacts
Yipeng Miao, Yenan Xu · 2024
Random Forest (RF) is a powerful machine learning algorithm for classification (RF-C) and regression (RF-R) problems and excels in feature importance analysis. It accurately evaluates the contribution of each feature in classification or regression, i.e., feature importance, by integrating multiple decision trees or regression trees. Previous studies have often utilized this property to help researchers understand which features are most critical to the prediction task. However, in this paper, this property is applied to the analysis of the variability of the impact of features on the aggregate, i.e., to observe whether there is a significant difference in the importance of the same feature in different aggregates; if there is a significant difference, the feature can be considered to have different impacts on different aggregates, and vice versa, the impacts can be considered to be the same. By combining the theories of feature importance and hypothesis testing in large-sample settings, this paper proposes an innovative method to realize the impact difference analysis, which is different from the traditional ANOVA. Finally, in order to verify the feasibility of the method, this study applies the above method on a real dataset and evaluates the feasibility of the final results in conjunction with one-way ANOVA. The experiment proved the feasibility of the method and demonstrated its effectiveness in analyzing the differences in the impact of characteristics on different aggregates.