Sample Reduction for SVMs via Data Structure Analysis
Defeng Wang, Daniel So Yeung, C. C. Tsang Eric · 2007
This paper presents a new sample reduction algorithm, sample reduction by data structure analysis (SR-DSA), for SVMs to improve their scalability. SR-DSA utilizes data structure information in determining which data points are not useful in learning the separating plane and could be removed. As this algorithm is performed before SVMs training, it avoids the problem suffered by most sample reduction methods whose choices of samples heavily depend on repeatedly training of SVMs. Experiments on both synthetic and real world datasets have shown that SR-DSA is capable of reducing the number of samples as well as the time for SVMs training while maintaining high testing accuracy.