Performance evaluation of distributed framework over YARN cluster manager

Solaimurugan Vellaipandiyan, Penmetsa V. Krishna Raja · 2016

The world is driven by data intensive applications and eventually volume of data is generated out of these applications. These data do not yield benefits in its native form. It requires a different mechanism to collect, store, process and derive insights from these data. Apache Hadoop is one of the open source frameworks intended for this purpose. Apache Hadoop provides mechanism to distribute the data across cluster and parallel processing on it. Apache Spark is another emerging in-memory framework for large-scale, faster data processing. Such system requires a lot of system resources like memory, CPU and storage depending on the data size. In this paper we empirically present comparative approach between Apache Hadoop and related tools (Apache Spark) for processing (performing a search operation on different datasets.) large amount of data. Consequently, various performance behaviors like CPU utilization, memory, RSS and VSZ were recorded along with time to execute the job on test cluster environment. We find that performance of the frameworks degrades based on the size, memory and time factor.

Read the paper · More papers on PaperTik