Improving performance of decision trees with multi-edit-nearest-neighbor algorithm

Nianyi Chen · Kongzhi yu juece · 2003

Noises and overlapped regions existing in training samples hurt the simplicity and generality of decision trees. To solve this problem, a sample selection algorithm based on multi-edit-nearest-neighbor rule is proposed. This algorithm, under ideal conditions, can eliminate the noise satisfying some prerequisites, purify the overlapped region according to its members′ posterior probabilities, and finally form a Bayesian boundary between samples of different classes. When applied to an appropriate trainingdataset,itobviouslycutsdownthesize of resulting decision trees without sacrificing the accuracy. This improves both the understandability and generality of decision trees.

Read the paper · More papers on PaperTik