Performance bounds and buffer space requirements for concurrent processors

William H. Mangione-Smith · Deep Blue (University of Michigan) · 1992

Scientific programs are typically characterized as floating-point intensive loop-dominated tasks with large amounts of exploitable parallelism. A wide range of concurrent processor architectures have been proposed to capitalize on this parallelism, e.g. vector, multi-threaded, superscalar, and VLIW. Even though these architectures all attempt to exploit the same performance potential, the wide range of hardware structures has to date discouraged a unified performance evaluation technique. This thesis develops and uses application-specific upper bounds on performance to evaluate a wide range of concurrent machines. The method focuses on the bandwidth of critical machine units, abstracting away many details of the underlying architectures. Three machines are evaluated as examples: the RISC DECstation 3100, the superscalar IBM RS/6000, and the Decoupled Access-Execute Astronautics ZS-1. When compared to measured performance, performance bounds reveal the overall efficiency of execution and the effective utilization of the critical machine units. These bounds have proven useful for evaluating specific machines, comparing two or more machines and improving the performance of applications codes on these computers. Achieving high performance on machines with long latency operations demands increased buffer space, usually in the form of registers. The dependence on buffer space for high performance is highlighted by the three example machines mentioned above, each of which uses a fundamentally different structure for its buffer space. A method is presented for calculating minimal buffer space requirements for an application running at its optimum performance as machine latencies are varied. This method allows an architect to examine fundamental consequences of various buffer space and latency design choices. One such issue concerns the number of registers required in a vector machine, versus their length. The results presented here suggest that the historic trend toward longer vector registers as latencies have increased is the wrong approach and that greater resources should be spent on increasing the number of available registers even if their length is greatly reduced.

Read the paper · More papers on PaperTik