Optimal thresholds for classification trees using non parametric predictive inference

Masad A. Alrasheedi, Tahani Coolen‐Maturi, Frank P. A. Coolen · Communication in Statistics- Theory and Methods · 2025

In data mining, classification trees are used to assign a new observation to one of a set of predefined classes based on the attributes of the observation. They are constructed recursively through a top-down approach using repeated splits of the training dataset, which is a subset of the full data. When the dataset includes continuous-valued attributes, it is necessary to select appropriate threshold values to determine the classes and split the data. In recent years, nonparametric predictive inference (NPI) has been introduced for selecting optimal thresholds for two- and three-class classification problems, where the inferences are explicitly based on a given number of future observations and target proportions. The NPI-based threshold selection method has previously been applied in the context of receiver operating characteristic (ROC) analysis, but not for building classification trees. Due to its predictive nature, the NPI-based threshold selection method is well suited to classification tree construction, as the primary goal of such trees is prediction. In this article, we present new classification algorithms for building classification trees using the NPI approach for selecting optimal thresholds. We introduce a new procedure for selecting the optimal target proportions by optimising classification performance on test data. Various measures are used to evaluate and compare the performance of the NPI2-Tree and NPI3-Tree classification algorithms with other methods from the literature. Experimental results show that our proposed algorithms perform well, achieving accuracies of 83.84% for the NPI2-Tree and 82.87% for the NPI3-Tree. These results suggest that the proposed classification algorithms may serve as viable alternatives to existing methods.

Read the paper · More papers on PaperTik