Characterizing Scheduling Delay for Low-Latency Data Analytics Workloads
Wei Chen, Aidi Pi, Shaoqi Wang, Xiaobo Zhou · 2018
Data analytics workloads are shifting to shorter task execution time, higher degree of parallelism, and execution on faster hardware. As a result, job scheduling is becoming a bottleneck, which needs to offer extreme low-latency, massive throughput, and high scalability. However, few efforts have been focused on systematically understanding the scheduling delay. In this paper, we propose a method and develop a tool, SD-checker, that decomposes the job scheduling delay into multiple components and characterizes each by extensive experiments. SDchecker extracts event messages through mining both cluster scheduler logs and application logs, and constructs a scheduling order graph for the ease of analysis. We use SDchecker to evaluate Spark-SQL on a popular cluster scheduler Yarn. Results show that the scheduling delay may account for 60% of the job runtime of small data analytics workloads. After decomposing the total scheduling delay, we find Spark itself contributes 70% of the delay. Through the evaluation and analysis, we conclude that (1) The causes of scheduling delay are determined by many factors, and (2) The job scheduling is not well optimized yet, and far from ideal for low-latency data analytics workloads.