The floating point performance of a superscalar SPARC processor

Roland L. Lee, Alex Y. Kwok, Fayé A. Briggs · 1991

The floating point performance of superscalar SPARC processors is evaluated based on empirical data horn 12 benchmarks.This evaluation is done in the context of two software instruction scheduling optimization, loop unrolling and software pipelining, and for three machine models we term, 1-scalar, 2-scalar and 4-scalar.We also consider the effect of the memory system on the performance improvements.Superscalar hardware alone exhibit little performance improvement without software optimization.Of the two scheduling methods we study, software pipelining more effectively takes advantage of increased hardware parallelism, and achieves near optimal speedup on the 4-scalar machine model.The performance of loop unrolling is restricted by the limited number of floating point registers in the SPARC architecture.The best performance level is obtained by applying both optimization techniques.A superscalar SPARC processor can provide improved tloating point performance but with significant software and hardware development costs.

Read the paper · More papers on PaperTik