Characterizing Task Usage Shapes in Google Compute Clusters

Qi Zhang, Joseph L. Hellerstein, Raouf Boutaba · 2011

The increase in scale and complexity of large compute clusters motivates a need for representative workload benchmarkstoevaluatetheperformanceimpactofsystemchanges, so as to assist in designing better scheduling algorithms and in carrying out management activities. To achieve this goal, it is necessary to construct workload characterizations from which realistic performance benchmarks can be created. In this paper, wefocus on characterizingrun-timetaskresource usage for CPU, memory and disk. The goal is to find an accurate characterization that can faithfully reproduce the performance of historical workload traces in terms of key performance metrics, such as task wait time and machine resource utilization. Through experiments using workload traces from Google production clusters, we find that simply using the mean of task usage can generate synthetic workload traces that accurately reproduce resource utilizations and task waiting time. This seemingly surprising result can be justified by the fact that resource usage for CPU, memory and disk are relatively stable over time for the majority of the tasks. Our work not only presents a simple technique for constructing realistic workload benchmarks, but also provides insights into understanding workload performance in production compute clusters. 1.

Read the paper · More papers on PaperTik