Optimized Parallelization of Binary Classification Algorithms Based on Spark
Yushui Geng, Jianguo Zhang · 2016
In recent years, parallel computing and machine learning have gradually became a research hotspot in big data field. Now users who are familiar with the R language can easily complete the data analysis on the platform of Apache Spark with the programming interface called SparkR provided by Apache Spark. There are several realized classification algorithms (including multiple classification and binary classification) in SparkR, in order to speed up the training process in binary-class classification, this paper proposes a theory of adopting an iterative calculation model for the local optimization on the basis of conventional parallelization. The experimental result shows that the design of parallelization of binary classification algorithm based on spark is much better and faster comparing with the MapReduce on Hadoop.