Improving performance of decision trees with multi-edit-nearest-neighbor algorithm
Nianyi Chen · Kongzhi yu juece · 2003
Noises and overlapped regions existing in training samples hurt the simplicity and generality of decision trees. To solve this problem, a sample selection algorithm based on multi-edit-nearest-neighbor rule is proposed. This algorithm, under ideal conditions, can eliminate the noise satisfying some prerequisites, purify the overlapped region according to its members′ posterior probabilities, and finally form a Bayesian boundary between samples of different classes. When applied to an appropriate trainingdataset,itobviouslycutsdownthesize of resulting decision trees without sacrificing the accuracy. This improves both the understandability and generality of decision trees.