Compile-time performance prediction of scientific programs
David Padua, Gheorghe Calin Cascaval · 2000
In this disertation we present a compile-time performance prediction environment for Fortran scientific programs. The performance data are expressed as symbolic expressions, with variables for program constructs, input data size and machine parameters. We focus on modeling the processor and its memory hierarchy. The results from the static estimation can be used to drive optimizations or can be displayed using performance visualization tools. The integration of our model within the Delphi system allows the user to do performance tuning and scalability analysis faster and easier than by using instrumentation. The main contribution of this work is the cache behavior estimation using the stack distances algorithm. We have designed and implemented a compile-time algorithm that computes the stack histogram at compile-time. We use the stack histogram to predict program performance statically with very good accuracy. Experimental results are presented for two processor/memory architectures, the MIPS R10000 and UltraSparc IIi. The most interesting feature of the stack algorithm is that once the histogram is computed, the number of cache misses can be estimated for any cache size. We use stack distances to quantify locality and we show that the average locality computed using stack distances is a very reliable metric. A new algorithm for stack processing, that is 30% faster than the best know algorithm on the suite of programs traced, is also presented.