A comparative between hadoop mapreduce and apache Spark on HDFS
Mohamed Saouabi, Abdellah Ezzati · 2017
Data is growing now in a very high speed with a large volume, Spark and MapReduce1 both provide a processing model for analyzing and managing this large data -Big Data- stored on HDFS. In this paper, we discuss a comparative between Apache Spark and Hadoop MapReduce using the machine learning algorithms, k-means and logistic regression. This comparative is done through two experiences, the first one using the same programming language java, and the second using different programming languages.