Parallelized incremental support vector machines based on MapReduce and Bagging technique

Jun Zhao, Zhu Liang, Yong Yang · 2012

One of the mainstream research fields in learning from empirical data by support vector machines (SVM) is an implementation of the incremental learning schemes when the training dataset is huge. Moreover, the challenge of applying incremental SVMs on huge data sets comes from the fact that the amount of computer memory and learning time required along with the amount of dataset increased. In this paper, a parallelized incremental SVM (PISVM) learning algorithm for huge data is proposed. The parallel programming model of MapReduce is introduced and combined with incremental learning method. Each individual SVM is independently trained based on the randomly selected training samples via bootstrap technique, and learns from the new samples independently also. The final decision is made according to the majority voting by all SVMs. Experiment results on UCI standard data sets show that the training time can be reduced and the accuracy can be ensured for the proposed algorithm.

Read the paper · More papers on PaperTik