Performance comparision of Hadoop and spark engine

Akaash Vishal Hazarika, G Jagadeesh Sai Raghu Ram, Eeti Jain · 2017 International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC) · 2017

Data has been growing at an exponential rate in the recent era. This data has to be processed and analyzed carefully to get new insights. So there is a need for a platform that can perform efficient data processing, which leads to the building up of platforms like Hadoop and Spark. Hadoop Map Reduce model helped to process it in a distributed fashion with computation being done at multiple nodes. In spite of the remarkable processing power it still, has some shortcomings. Spark, a computation engine can solve some of the problems involving Iterative/Machine learning queries by caching some of the results from previous queries. Although Spark is faster than Hadoop in most of the iterative applications it is constrained by memory requirements. This paper briefly discusses Spark and Hadoop architecture, their theoretical differences and the comparison of their performance.

Read the paper · More papers on PaperTik