A Comparative Study of Big Data Processing: Hadoop vs. Spark
Meghna Sharma, Jagdeep Kaur · International Conference on Computing for Sustainable Global Development · 2019
Apache Spark and Hadoop’s MapReduce are two very important tools used for Big Data processing. The processing started with Hadoop’s MapReduce Framework but suffers from many disadvantages due to multiple disc processing operations. The drawbacks of the traditional big data processing have been overcome by in memory handling framework like Spark. In some aspects they go hand in hand as due to lack of file system in Spark, it needs to depend upon MapReduce. This paper has shown the extensive study on various tools related to Big Data processing and has done extensive comparison on MapReduce Vs Spark. The frameworks have been studied on real time datasets and finally compared in terms of processing time. Spark showing the remarkable improvement over MapReduce.