DISTRIBUTION BASED TREES ARE MORE ACCURATE

Nong Shang, Leo Breiman · 2013

Classification trees are attractive in that they present a simple and easily understandable structure. But on many data sets their accuracy is far from optimal. Much of this lack of accuracy is due to their instability--small changes in the data can lead to large changes in the resulting tree. This instability is the reason that combining many trees by voting can lead to dramatic decreases in test set error (Breiman[1995]). But combining trees loses the simple structure. To keep the simple structure and improve accuracy, a way must be found to reduce the instability in the construction. If we knew the true probability distribution of the inputs and outputs, then the splits in the tree could be based on this distribution and give more accuracy then the splits based on a finite data set. So we turn the tree procedure around--instead of basing the splits on the data, the data is used to estimate the input-output probability distribution and the splits are then based on this estimate. We g...

Read the paper · More papers on PaperTik