Efficient Real-time Earliest Deadline First based scheduling for Apache Spark

Laurentiu-Florin Neciu, Florin Pop, Elena‐Simona Apostol, Ciprian‐Octavian Truică · 2021

Apache Spark is a distributed computing framework for fast in-memory data analysis, Machine Learning jobs, and SQL queries that employs Resilient Distributed Dataset for distributing data and Directed Acyclic Graph for scheduling computations. Currently, Spark provides two scheduling policies for tasks pending execution: a First In First Out policy and a FAIR policy, providing no support for deadline-based real-time scheduling. In this paper, we present a new system designed to accept deadlines for heterogeneous Spark Jobs and perform a real-time scheduling policy based on Earliest Deadline First (EDF). To showcase the efficiency of our scheduling policy, we compare and analyze the performance of our solution with the current Spark execution policies in terms of job lateness. We empirically prove that our real-time policy provides much lower lateness given suitable constraints.

Read the paper · More papers on PaperTik