Stabilization Of Regression Trees

T. Urban, Thomas Kämpke · WIT transactions on information and communication technologies · 2000

In this paper, we present a hierarchical approach to simultaneous regression and classification. Regression obviously becomes more accurate by assessing a regression surface to each of a given class of a finite sample set compared to one regression surface for the sample as a whole. For class formation, a tree of regression surfaces is constructed that balances minimization of the regression error and towards unseen data. Common tree-structured regression algorithms split nodes according to an independent variable. Terminal nodes correspond to one specific class. Such constructions gave rise to less complicated approaches that have mainly been used for adaptive classification in machine learning. The considerable advantage of regression trees over a single regression is often set off by trees' poor behaviour on unseen data of supposedly the same nature. Also, a tree formation solely based on minimizing the regression error may not generate information by which to assign new data to terminal nodes. Generalization is addressed by a stabilization operation. Each sample vector is assigned to its nearest neighbour which in turn is assigned to its nearest neighbour etc. This results in so-called neighbour chains that partition the sample set. Class splitting during tree formation is here restricted to classes that contain neighbour chains completely. For classification, unseen sample data are assigned to its closest neighbour used in tree formation. The regression surface of the class containing the corresponding neighbour chain is then used for estimation. Data Mining II, C.A. Brebbia & N.F.F. Ebecken (Editors) © 2000 WIT Press, www.witpress.com, ISBN 1-85312-821-X

Read the paper · More papers on PaperTik