An Improved Error-Based Pruning Algorithm of Decision Trees on Large Data Sets

Yi Peng, Yutong Lu, Zhiguang Chen · 2021

Decision tree is one of the models that are often used in classification. Pruning is necessary for decision tree in order to prevent overfitting. With the advent of the era of big data, there is one well-known folklore that, using decision tree pruning algorithms, growing trees with increasingly larger amounts of training data will result in larger tree size even when accuracy does not increase. This paper discusses issues in pruning a scalable classifier and presents the design of a new parallel algorithm based on error-based pruning. The improved error-based pruning (IEBP) uses k-fold cross validation and t-test to choose the optimal value for the certainty factor, which controls the pruning, instead of using the default one. Different experiments have been made to compare the behavior of the behavior of IEBP with the behavior of error-based pruning and reduced-error pruning. The Experimental results support the conclusion that varying the certainty factor allows significantly smaller trees to be obtained with no or minimal accuracy loss on large data sets.

Read the paper · More papers on PaperTik