Approximate Splitting for Ensembles of Trees using Histograms

Chandrika Kamath, Erick Cantú‐Paz, David Littau · 2002

1 Introduction Ensembles of classifiers have become an active topic of research in the data mining community. They not only provide a simple way of improving the accuracy of the classifier [3, 16, 24, 2], but also have the potential for on-line classification of large databases that do not fit into memory [4]. In addition, some approaches to the generation of ensembles can be easily parallelized, enabling a reduction in the time taken to create the classifier on a multiprocessor system [17]. There are several different ways in which ensembles can be generated and the resulting output combined to classify new instances. Implicit in many of these ensembles is the concept of randomness that is introduced either through the randomization of the training set, or the randomization of the classifier itself.

Read the paper · More papers on PaperTik