Reduced input data sets selection for SPEC CPUint2006
Virginia Escuder, Rafael García Rico · 2009
SPEC CPU benchmark suite has become the most frequently used suite for computer architecture research. The workload is designed to stress the hardware of the machines for the next generation. Consequently, the executed instruction count has been considerably increased in the SPEC CPU2006, compared to the previous suites. But this fact results in prohibitive experimentation time or resources requirements for research when using simulation techniques or working in embedded systems development. In the Literature has been described a wide range of approaches for reducing the workloads in experimental environments. In this work an extensive revision of them is offered. Moreover, a detailed analysis of the influence of input data sets in the workload of the CPUint2006 suite is presented. As a result, a suggestion of alternative workload with 2 orders of magnitude dynamic executed instructions count lower than test workload is proposed and characterized. The aim of our work is to help researchers finding a representative set of workloads for the SPEC CPUint2006 programs to use in their experiments whenever they have to discard using the reference workload due to time or resource constraints. Reduced input data sets selection for SPEC CPUint2006 3 1. SPEC benchmarks SPEC is the acronym for Standard Performance Evaluation Corporation, a non-profit organization whose purpose is to define and maintain a set of standard benchmarks for computer systems and make them available to the users of such systems as a common reference point in the evaluation of computer performance. It is participated by computer manufacturers, system integrators, consultants, publishers, universities and research organizations [50]. Since its foundation in 1988, the SPEC consortium has developed and distributed technically reliable benchmarks based on real applications. The selection of inputs and workload is performed by the consensus amongst consortium members willing for a transparent, comparable, reproducible and nonproprietary solution, as the organization aim is “an ounce of honest data is worth a pound of marketing hype”. SPEC currently offers testbench for the evaluation of different aspects of computation such as performance of CPU, graphics, distributed Java computing, web servers, and network file systems. 1.1. The SPEC CPU suite The present technical report belongs to the field of performance quantification for intensivecomputing. The testbench set from the SPEC organization that best fits this field is the SPEC CPU suite. The SPEC CPU suite aim is to be representative of programming style and application fields of real worldwide computer-intensive workload. The first delivered set in 1989 had 10 programs and it was known as SPECmark. The most recent generation of the set is from 2006 (SPEC CPU2006) and it is made of 29 programs classified into two groups: 12 programs for integer computation (SPEC CPUint2006) and 17 programs for floating point computation (SPEC CPUfp2006) [20]. A more complete historical perspective of computer-intensive tests can be found in Henning’s work [23]. SPEC CPU programs are well known real world applications written in high level, portable, language (C or C++) with slight code modifications in order to minimize input/output and thus let the processor, memory and compiler be the factors under evaluation. In fact, it is a requirement that input/output workload is less than 5% and the article from Ye, Ray and Kaeli show that the I/O activity of the SPEC CPU2006 is far less than this limit [58]. Another requirement is that memory consumption and execution time should be significant for each generation of computers with growing power and capacity. The organization keeps tight restrictions of evaluation rules affecting code, compilation flags and other aspect of execution environment of the test programs. The workload sets the amount of processing performed by each benchmark run. The tools distributed by SPEC allow the specification of three different sizes of input data producing different workloads: test, train and reference (“ref” for short). The reference size stands for the reference workload, that is, the input data and command-line options when applicable, used for actual measurements. The test input sets are only used to check that programs compile and execute correctly before launching a real run or to tune optimization options. Similarly, the train inputs are used for profile-based compiler optimizations, so the reference set is the only reportable set. The SPEC suite has a widespread usage by computer vendors, it is widely accepted by consumers and it is very commonly found too in the academic and research worlds although there is an outstanding debate about how to use it, its convenience and drawbacks, and whether is it or not necessary to design an alternative for research usage, etc. The goal of this present work is reduced to the integer benchmarks (SPEC CPUint2006). A thoroughly characterization of SPEC CPUint2006 can be found in the technical report “SPEC CPUint2006 characterization” [13]. 1.2. Controversia in the academic world There is a live debate in the academic world lasting several years about whether the SPEC tests are adequate or not for research. As we said before, the workload has been design to contrast the hardware of the machines for the next following years. Consequently, the executed instructions count has considerably increased as well as the size of memory map used. Then, one of the problems stated is that they are far too large (in size and execution time) for experimentation and analysis. This has lead to some misuse by the research community in the attempt to cut down the size of experiments, according to some authors who