Multiple submodels parallel support vector machine on spark
Chang Liu, Bin Wu, Yi Yang, Zhihong Guo · 2016
The Support Vector Machine (SVM) is a classical classification algorithm that has a wide range of application. With kernel function, SVM can dispose the datasets that are not linearly separable in their original feature space, making it more flexible in practical use compared with linear model. However, its complexity in training is an obstacle to large-scale dataset handling. This paper proposes the Multiple Submodels Parallel SVM (MSM-SvM) on Spark to accelerate the training of nonlinear SVM with computer cluster. Making use of SVM' s model theory, we introduce the data-splitting method based on clustering to enable the parallel training process and approximate global solution with a number of local submodels. We also cover the multi-classification with “one-against-one” strategy. The experiment shows that MSM-SVM do well in not only binary classification but also the multiclass one. It outperforms the SVM With MiniBatch SGD tool in Spark MLlib in predicting accuracy for binary case. And for almost all our experiment datasets, MSM-SVM gives similar predicting performance with LIBSVM while costing much less time.