the Reduction of Irrelevant Training Data
Robert S. Lynch, Peter Willett · 1998
In this paper, performance of a previously introduced method of data reduction (referred to as the Bayesian Data Reduction Algorithm) is demonstrated which uses a noninformative (i.e., Dirichlet distribution) prior on the symbol probabilities. The algorithm employs a [‘greedy” approach that relies on the average conditional probability of error as a metric for making data reducing decisions. Performance is compared to a neural network at classifying discrete feature vectors containing binary and ternary valued features, and it is shown that the Bayesian Data Reduction Algorithm is superior. However, performance of both schemes is also shown to degrade as the quantization fineness is increased with ternary valued features.