Support Vector Machine Usingk-Spatial Medians Clustering and Recovery Process
Sungwan Bang, Ja‐Yong Koo, Myoungshic Jhun · Communications in Statistics - Simulation and Computation · 2010
Even though support vector machine has been successfully applied to various classification problems with its flexibility and high classification accuracy, it is not suitable for classification of large data sets because its computational complexity grows rapidly as the size of data set increases. SVM using k-means clustering (KM-SVM) is a fast algorithm which has been developed to accelerate both computation and prediction of SVM classifiers. However, it seems likely that the data set is contaminated by outliers in real-world situations, and k-means clustering is sensitive to these outliers. Therefore, we propose to combine k-spatial medians clustering with SVM (KS-SVM) since k-spatial medians clustering is robust for outliers. In order to improve the classification accuracy in KS-SVM, furthermore, a recovery process based on KS-SVM (RKS-SVM) is also proposed in this article. Experiments show that KS-SVM can improve the performance of KM-SVM in terms of classification accuracy and number of support vector. It is also shown that the classification accuracy of RKS-SVM is better than one of KS-SVM, but there is a tradeoff between classification accuracy and the number of support vectors.