Dynamic Job Scheduling for Energy Savings in YARN: A Comparative Study With Benchmark Workloads

Vaibhav Pandey, Poonam Saini, Mansi Gupta, Aditya Bhardwaj · Concurrency and Computation Practice and Experience · 2025

ABSTRACT The energy efficiency of MapReduce systems has become a crucial aspect of managing big data analytics infrastructure. Hadoop is the most popular open‐source implementation of the MapReduce paradigm. Its scheduler plays a key role in ensuring this efficiency. To minimize energy consumption in slot‐based Hadoop systems, various MapReduce scheduling techniques have been developed. However, YARN‐based Hadoop schedulers have not been given much attention in terms of improving energy efficiency. Thus, addressing the MapReduce scheduling problem for YARN‐based Hadoop is essential for promoting sustainable cloud computing. In this paper, we propose a framework for the online MapReduce scheduling problem, where jobs arrive randomly and their arrival times are unknown. To tackle this issue, we developed two online scheduling algorithms that greedily select the job with the lowest energy usage at each iteration. We evaluate the proposed methods for large‐scale, three separate workloads of standard benchmark jobs, namely, WordCount (CPU‐bound), TeraSort (I/O‐bound), and NutchIndexing (mix‐bound). We also evaluate the performance for a mixed workload comprising all three benchmark jobs. The experimental results show that the proposed method considerably minimizes energy consumption for all workloads compared to the existing FAIR and capacity schedulers. Particularly, we find it to be 35% more energy‐efficient than the FAIR scheduler, contributing to sustainable cloud computing.

Read the paper · More papers on PaperTik