STABILIZER: Enabling Statistically Rigorous Performance Evaluation
Charlie Curtsinger, Emery D. Berger · 2012
Modern architectures have made program behavior brittle and unpredictable, making software performance highly dependent on its execution environment. Even apparently innocuous changes, such as changing the size of an unused environment variable, can—by altering memory layout and alignment—alter performance by 33 % to 300%. This unpredictability makes it difficult for programmers to debug or understand application performance. It also greatly complicates the evaluation of performance optimizations, since slight changes in the execution environment can have a greater impact on performance than a typical optimization. We present STABILIZER, a compiler and runtime system that enables statistically rigorous performance evaluation. STABILIZER eliminates measurement bias by comprehensively and repeatedly randomizing the placement of functions, stack frames, and heap objects in memory. Random placement makes anomalous layouts unlikely and independent of the environment, and re-randomization ensures they are shortlived when they do occur. We demonstrate that applications compiled with STABILIZER deliver normally-distributed execution times, enabling the use of standard statistical tools for hypothesis testing. We demonstrate its use by testing the effectiveness of standard optimizations used in the LLVM compiler; we find that, across the SPEC CPU2000 and CPU2006 benchmark suites, the effect of the-O3 optimization level versus-O2 is indistinguishable from noise. 1.