Optimized Parallelization of Binary Classification Algorithms Based on Spark

Yushui Geng, Jianguo Zhang · 2016

In recent years, parallel computing and machine learning have gradually became a research hotspot in big data field. Now users who are familiar with the R language can easily complete the data analysis on the platform of Apache Spark with the programming interface called SparkR provided by Apache Spark. There are several realized classification algorithms (including multiple classification and binary classification) in SparkR, in order to speed up the training process in binary-class classification, this paper proposes a theory of adopting an iterative calculation model for the local optimization on the basis of conventional parallelization. The experimental result shows that the design of parallelization of binary classification algorithm based on spark is much better and faster comparing with the MapReduce on Hadoop.

Read the paper · More papers on PaperTik