Impact of Memory Size on Bigdata Processing based on Hadoop and Spark

Seunghye Han, Wonseok Choi, Rayan Muwafiq, Yunmook Nah · 2017

Hadoop and Spark are well-known big data processing platforms. The main technologies of Hadoop are Hadoop Distributed File System and MapReduce processing. Hadoop stores intermediary data on Hadoop Distributed File System, which is a disk-based distributed file system, while Spark stores intermediary data in the memories of distributed computing nodes as Resilient Distributed Dataset. In this paper, we show how memory size affects distributed processing of large volume of data, by comparing the running time of K-means algorithm of HiBench benchmark on Hadoop and Spark clusters, with different size of memories allocated to data nodes. Our results show that Spark cluster is faster than Hadoop cluster as long as the memory size is big enough for the data size. But, with the increase of the data size, Hadoop cluster outperforms Spark cluster. When data size is bigger than memory cache, Spark has to replace disk data with memory cached data, and this situation causes performance degradation.

Read the paper · More papers on PaperTik