Why Many Benchmarks Might Be Compromised

Johannes Manner, Guido Wirtz · 2021

Benchmarking experiments often draw strong conclusions but lack information about the environmental influences like the hardware used to deploy the investigated system. Fairness and repeatability of these benchmarks are at least questionable. Developing for or migrating applications to the cloud or DevOps environments often requires performance testing, either for ensuring quality-of-service or for choosing the correct service parameters when deciding for a cloud offering. While building a benchmarking pipeline for cloud functions, the typical assumption is that a CPU scales the resources linearly to the used utilization. Due to heat generation, noise and other constraints, this is not the case due to the trade off between efficiency and performance. To investigate this trade off and its implications, we set up some experiments in order to evaluate the influence of these factors for benchmark results. We solely focus on Intel CPUs. Beginning with the second generation (Sandy Bridge), Intel uses their own scaling driver intel_pstate. Our results show that different settings for this scaling driver have a significant impact on the measured performance and therefore on the linear regression models we computed using LINPACK benchmarks. These benchmarks are executed at different CPU utilization points. An active intel_pstate scaling driver with enabled turbo boost and powersave governor reached a R2 of 0.7349, whereas the performance governor shows a significantly better, ideal determination coefficient with 0.9999 on a machine used in the benchmarks. Therefore, we propose a methodology for system calibration to ensure fair and repeatable benchmarks.

Read the paper · More papers on PaperTik