STUDY OF SCHEDULING TECHNIQUES IN HADOOP- MAPREDUCE

Devika Rankhambe, Sharmishtha Desai, Sarika Patil · 2013

Hadoop was designed mainly for running large batch jobs such as web indexing and log mining. Users submitted jobs to a queue and the cluster ran them in order. However, as organizations placed more data in their Hadoop clusters and developed more computations they wanted to run, sharing a MapReduce cluster between multiple users became important. Because of sharing, as all the data is in one place, users can run queries for the data they want to fetch and execute. To share a MapReduce cluster, support from the Hadoop job scheduler is needed to provide a guaranteed capacity for production jobs viz, load data, compute statistics, detect spam and good response time to interactive jobs while allocating resources fairly between users. Now, scheduler in Hadoop became a pluggable component, it has opened the door for innovation. Two schedulers were developed for multi-user workloads: the Fair Scheduler, developed at Facebook, and the Capacity Scheduler, developed at Yahoo. In this paper, we study various hadoop schedulers and current research improvements in those schedulers. We studied scheduling frameworks available to improve performance characteristics of Mapreduce jobs.

Read the paper · More papers on PaperTik