Concepts for Probabilistic and Possibilistic Induction of Decision Trees on Real World Data

Christian Borgelt, Jörg Gebhardt · 2004

The induction of decision trees from data is a well-known method for learning classifiers. The success of this method depends to a high degree on the measure used to select the next attribute, which, if tested, will improve the accuracy of the classification. This paper examines some possibility-based selection measures and compares them to probability- and information-based measures on real world datasets. The results show that possibility-based measures do not much worse with regard to classification accuracy, in certain cases they seem to do even slightly better. 1 Introduction An often used method to induce decision trees from data is a greedy algorithm, which inspects the conditional distribution of the considered classes within the given dataset for each attribute. These conditional distributions are valuated with some selection measure and the attribute yielding the highest (or lowest, depending on the measure) value is chosen to partition the set of cases. Then this procedure ...

Read the paper · More papers on PaperTik