Not Too Hot, Not Too Cold: The Bundled-SVM is Just Right!

Lawrence Shih, Yu-Han Chang, Jason D. M. Rennie · 2002

The Support Vector Machine (SVM) typi-cally outperforms other algorithms on text classification problems, but requires train-ing time roughly quadratic in the number of training documents. In contrast, linear time algorithms like Naive Bayes have lower per-formance, but can easily handle huge training sets. In this paper, we describe a technique that creates a continuum of classifiers be-tween the SVM and a Naive Bayes like algo-rithm. Included in that continuum is a clas-sifier that approximates SVM performance with linear training time. Another classifier on this continuum can outperform the SVM, yielding a breakeven point that beats other published results on Reuters-21578. We give empirical and theoretical evidence that our hybrid approach successfully navigates the tradeoffs between speed and performance. 1. Introduction & Related

Read the paper · More papers on PaperTik