SPLITTING METHODS FOR DECISION TREE INDUCTION:A COMPARISON OF TWO FAMILIES

Kweku-Muata Osei-Bryson, Kendall E. Giles · 2002

Decision tree (DT) induction is among the more popular of the data mining techniques. An important component of DT induction algorithms is the splitting method, with the most commonly used method being based on the conditional entropy family. However, it is well known that there is no single splitting method that will give the best performance for all problem instances. In this paper we explore the relative performance Conditional Entropy family and another family that is based on the Class-Attribute Mutual Information (CAMI) measure. Our results suggest that while some datasets are insensitive to the choice of splitting methods, other datasets are very sensitive to the choice of splitting methods. For example, some of the CAMI family methods may be more appropriate than GainRatio (GR) for datasets where all non-class attributes are nominal; some of the CAMI methods perform as well as GR for datasets where all the non-class attributes are either integer or continuous. Given the fact that it is never known beforehand which splitting method will lead to the best DT for the given dataset, and given the relatively good performance of the CAMI methods, it seems appropriate to suggest that splitting methods from the CAMI family should be included in data mining toolsets.

Read the paper · More papers on PaperTik