Evaluation and Design of Benchmark Suites
Jozo J. Dujmovic · 2024
Benchmark suites are most frequently designed for industrial evaluation of competitive computer systems and networks. Examples of such benchmark suites include SPEC 1 , TPC 2 , GPC 3 , PERFECT Club Benchmarks 4 , AIM benchmarks 5 , and others. In addition to benchmark suites sponsored by consortia of computer industry there are various collections of benchmarks designed by research organizations, companies, computer magazines, and individuals for benchmarking specific hardware and software systems. Examples of such benchmark suites include database benchmarks 6 , supercomputer benchmarks (Livermore loops 7 , NAS Parallel Benchmarks 8 , Lisp benchmarks 9 , Prolog benchmarks 10 , and many others 11 . In the majority of cases benchmark workloads are selected from a specific set of frequently used real workloads. The selection process is usually aimed at simultaneous satisfaction of two goals: (1) benchmark workloads should be a good functional representative of a given universe of real workloads, and (2) benchmark workloads should yield the same distribution of the utilization of system resources as real workloads. Due to the absence of specific quantitar tive methods for benchmark suite design, these goals primarily serve as guidelines in an intuitive workload selection process. Such a process cannot include proofs of the extent to which the goals are satisfied, and frequently yields relatively low reliability of benchmark results and excessive cost of benchmarking.