Evaluation and Analysis of Capacity Scheduler and Fair Scheduler in Hadoop Framework on Big Data Technology

Muhammad Salman, Diyanatul Husna, Adhitya Wicaksono, Anak Agung Putri Ratna · 2018

Apache Hadoop is an open source framework that implements MapReduce. It is scalable, reliable, and fault tolerant. Scheduling is an important process in Hadoop MapReduce. It is because scheduling has responsibility to allocate resources for running applications based on resource capacity, queues, running tasks, and the number of users. Changing single node to multi node Hadoop cluster can optimize HDFS, but quite costly. Scheduler performs the function of scheduling based on resource requirements, such as memory, CPU, disk, and network. The most general purpose of scheduling algorithm is minimizing the time of completing a task. Hadoop Scheduling is an independent module where users are able to design their own scheduler based on the application's actual need. So it can fulfill the specific need of the business in accordance with the desired result. This research will analyze the characteristic of Capacity Scheduler and Fair Scheduler.

Read the paper · More papers on PaperTik