Cross-layer Scheduling for MapReduce-based Big Data Workflows in Heterogeneous Hadoop Systems

Yijie Zhang, Chase Qishi Wu, Aiqin Hou · 2025

The performance of big data workflows depends on both the workflow mapping scheme, which determines task assignment and container allocation in Hadoop, and the on-node scheduling policy, which governs resource allocation and container provisioning. Most research on big data workflow scheduling focuses solely on workflow mapping, achieving only limited success. We conduct an in-depth investigation into the impact of node-level scheduling on overall workflow performance and explore the benefits of combining these two levels of scheduling (workflow- and node-level). We formulate a generic problem that considers cross-layer scheduling to minimize the end-to-end delay of MapReduce-based big data workflows in the Hadoop system. The efficacy of our proposed solution, compared with existing methods, is demonstrated through extensive simulations and proof-of-concept experiments using real-life big data workflows deployed on a real-life cluster.

Read the paper · More papers on PaperTik