YARN versus MapReduce — A comparative study
Sarah Shaikh, Deepali Rahul Vora · International Conference on Computing for Sustainable Global Development · 2016
The amount of digital data has exploded in recent years. One of the most popular tools for storage and processing of such huge digital data is Apache Hadoop[8]. The initial version of Hadoop i.e Hadoop 1.x supports distributed processing of large scale data using core MapReduce processing engine. MapReduce version l(MRvl) allows processing of big amount of data(i.e BigData) in-parallel on large clusters of computers. Even though Hadoop can be considered as a fast, scalable, cost-effective and fault-tolerant solution to big data problem, certain limitations exists in the Hadoop 1.x. However, things changed with the introduction of Hadoop 2.x. Hadoop 2.0 uses YARN(Yet Another Resource Negotiator), which separates the resource management and job scheduling task. This paper discusses the advantages YARN offered over the previous version of processing framework in Hadoop and also compares MapReduce and YARN based on some selected parameters.