Confusion by All Means
Muhammad Iqbal, Lizy K. John · 2006
Abstract—Performance of computers is usually measured by using benchmark suites. There has been a long debate among computer architects on how to aggregate the individual program results to present a summary of performance over the entire suite. Many researchers have criticised the use of Geometric Mean (GM) but SPEC continues to use it to report performance. Mashey [7] has strongly supported the use of GM. According to Mashey, the programs in a benchmark suite like SPEC are samples of some population of programs. It is important that we model the distribution of population correctly before calculating any statisitics and making conclusions based on those statisitcs. Mashey also conjectures that Lognormal distribution is a better model than the Normal ditribution for such benchmark suites. Since GM is the back-transformed average of a Lognormal distribution, its use as a measure of central tendency is statistically correct. In this study, we evaluate the correctness of this Lognormal assumption using the large repository of performance results for SPEC CPU2006 published on SPEC’s website. Utilizing different tests for normality, we find out that although Lognormal distribution models the performance results better than the Normal distribution, there is a very large percentage of machines which are neither Normal nor Lognormal. Our study indicates that most of the non-normality is caused by small number of outliers. We study the causes of these outliers and evaluate the use of Coefficeint-of-Variance to identify outliers. We also present some suggestions on how to deal with these outliers.