Optimal multi-splitting of numeric ranges for decision tree induction

Pen Lutu · Unisa Institutional Repository (University of South Africa) · 2001

Data mining is the process of extracting informative patterns from data stored in a database or data warehouse. Decision tree induction algorithms, from the area of machine learning are well suited for building classification models in data mining. The handling of continuous-valued attributes in decision tree induction has received a lot of research attention in recent years. Typically, an evaluation function is used to dynamically select the best multi-split for the range of values of a continuous-valued attribute. This paper discusses useful and well behaved evaluation functions and proposes an algorithm for optimal multi-splitting.

Read the paper · More papers on PaperTik