Classifying one billion data with a new distributed svm algorithm
Thanh‐Nghi Do, François Poulet · 2006
The new incremental, parallel and distributed Support Vector Machine (SVM) algorithm using linear or non linear kernels proposed in this paper aims at classifying very large datasets on standard personal computers (PCs). SVM and kernel related methods have shown to build accurate models but the learning task usually needs a quadratic program so that the learning task for large datasets requires large memory capacity and long time. We extend a recent Least Squares SVM (LS-SVM) proposed by Suykens and Vandewalle for building incremental, parallel and distributed SVM algorithm. The new algorithm is very fast and can handle very large datasets in linear or non- linear classification tasks on PCs. An example of the effectiveness is given with the linear classification into two classes of one billion datapoints in 20-dimensional input space in some minutes on ten PCs (3 GHz Pentium IV, 512 MB RAM).