Dynamic deadline-constraint scheduler for Hadoop YARN
Xiaofei Hou, Tarun Kumar, Johnson P. Thomas, Hong Liu · 2017
Hadoop is one of the most popular distributed platforms for big data. Apache YARN is the second generation of MapReduce. YARN has three built-in schedulers: the FIFO, Fair and Capacity Scheduler. Though these schedulers provide users different methods to allocate resources of a Hadoop cluster to execute their MapReduce jobs, they do not guarantee that their jobs will be executed within a specific deadline. In this paper, we propose a deadline constraint scheduler algorithm for Hadoop. This algorithm uses a statistical approach to measure the performance of data nodes and based on this information the proposed algorithm creates several checkpoints to monitor the progress of a job. Based on the progress of jobs at every checkpoint the proposed scheduler will assign them to different job queues. These queues will have different priorities and the proportion of resources used by these queues will depend on their priority. The results of our experiments show that the proposed scheduler ensures that jobs will be completed within a given deadline whereas the native schedulers cannot guarantee this. Moreover, the average job execution time in the proposed scheduler is less when compared to the other schedulers.