Enhancing the Explainability of Gradient Boosting Classification Through Comparable Samples Selection

Emilien Boizard, Gilles Chardon, Frédéric P. Pascal · 2024

Gradient-Boosted Decision Trees (GBDT) stand out as a powerful Machine Learning tool in tackling classification and regression tasks. Despite its effectiveness, GBDT, like other ensemble methods, suffers from a lack of explainability. Understanding these models' inner workings is crucial for comprehensively grasping their decision-making processes. In this study, we propose a method to enhance the explainability of GBDT, focusing on identifying specific training data points termed “comparable samples,” which play a pivotal role in the model's predictions. Inspired by the Frank-Wolfe algorithm, we introduce Explainable Gradient Boosting (ExpGB), which aims to shed light on the relationships between input data attributes and model predictions. ExpGB operates by ranking training samples based on their decomposition coefficients within the algorithm's output. Higher weights assigned to particular training samples indicate a closer resemblance to the sample being analyzed. To validate the efficiency of our approach, we conduct a comparative analysis across three diverse datasets, contrasting ExpGB with traditional GBDT algorithms. Through this analysis, we evaluate the quality of estimation provided by ExpGB, thereby enhancing our understanding of GBDT's workings.

Read the paper · More papers on PaperTik