Exploring Support Vector Machines for Big Data Analyses

Siyang Lu, Yihong Chen, Xiaolin Zhu, Ziyi Wang, Yangjun Ou, Yuhang Xie · 2021

The traditional support vector machines perform well in classification and prediction on small and medium-sized data sets, but there are some problems such as the low training efficiency and the low accuracy in large sample number, high dimension and large-scale data sets. Meanwhile, with the rise of distributed computing platforms such as the Spark suitable for big data analyses, more and more scholars at home and abroad turn their research direction to the distributed machine learning algorithms Therefore, in order to carry out the research on support vector machine for big data analyses, this paper explores the related researches and current situations of support vector machine, including: in-depth analysis of the algorithm principle of support vector machines, systematical investigation of the improved methods of support vector machines for the big data analyses, and distributed support vector machines under the Spark platform. Then, combined with the parallelization mechanism of the Spark, the some future research directions of support vector machine are investigated: for optimizing the accuracy of training results, some special matrix calculation skills should be added; In term of the research on SVM under the Spark platform, some better optimization methods from the perspective of dimension and partition can be found.

Read the paper · More papers on PaperTik