Text bundling: statistics-based data reduction
Lawrence Shih, Jason D. M. Rennie, Yu-Han Chang, David R. Karger · 2003
As text corpora become larger, tradeoffs be-tween speed and accuracy become critical: slow but accurate methods may not complete in a practical amount of time. In order to make the training data a manageable size, a data reduction technique may be necessary. Subsampling, for example, speeds up a classi-fier by randomly removing training points. In this paper, we describe an alternate method for reducing the number of training points by combining training points such that impor-tant statistical information is retained. Our algorithm keeps the same statistics that fast, linear-time text algorithms like Rocchio and Naive Bayes use. We provide empirical re-sults that show our data reduction technique compares favorably to three other data re-duction techniques on four standard text cor-pora. 1.