Work-in-Progress: Towards Efficient and Scalable Big Data Analytics: Mapreduce vs. RDD’s

Sheetanshu Srivastava, Akriti Nigam, Rashmi Kumari · 2017

Large-scale data analytics in the recent years has gained prominence due to explosion in data generated from the introduction of varied digital sources such as the internet, social media etc. With the ever increasing size of data and introduction of machine learning methodologies to learn consecutively from this data, distributed programming models such as MapReduce and its open source implementation Hadoop are facing performance issues. The present work discusses the concept of RDDs (Resilient distributed datasets) and it's open source implementation: Apache Spark. Hadoop and MapReduce pay a significant cost for reloading same data from the disk thus wasting significant input/output cycle as well as network bandwidth. Comparisons have been shown in terms of memory occupancy, execution time and lines of codes required in both the paradigms.

Read the paper · More papers on PaperTik