Instruction-level characterization of the Cray Y-MP processor
Sriram Vajapeyam · Minds at UW (University of Wisconsin) · 1992
Evolutionary computer architecture design fundamentally relies on empirical knowledge of workload characteristics and of dynamic program usage of machine features. While vector machines have thus far dominated supercomputing, detailed empirical studies of these machines have not been reported to date in the literature. We report an instruction-level study of a CRAY Y-MP vector processor, using as benchmarks the PERFECT Club suite of scientific and engineering applications. We study benchmarks that are automatically optimized and vectorized by the Cray Research, Inc. state-of-the-art production FORTRAN compiler. Several easily implementable manual optimizations provide significant performance improvements today over the best efforts of a compiler. Therefore, we also study a version of the benchmarks, hand-optimized by a Cray Research, Inc. team, that won the 1990 Gordon Bell-PERFECT award for the fastest version of the PERFECT Suite on any machine. We study only user routines in both cases. We observe that the user routines' vectorization level varies widely for the compiler-optimized version of the programs. While hand optimizations do improve the vectorization level, several benchmarks still have vectorization levels below 80%. Consequently, the non-vector features of vector machines are important to program performance. The scalar code in the programs contains numerous address calculation operations and Cray-specific miscellaneous operations, both important to performance. Scalar code's utilization of functional unit pipeline parallelism is low; latency improvements resulting from decreased pipelining could enhance scalar performance. Towards exploiting data dependencies for faster scalar code execution, we also characterize the dependencies within scalar basic blocks. Large basic blocks are significant in number as well as important to program performance; compiler optimization techniques geared towards large basic blocks are desirable. The vectorization level of memory operations is usually very high, thus emphasizing the need for high memory bandwidth. For the CRAY Y-MP hardware and for techniques currently employed by the Cray Research compiler, the peak issue rate of one instruction per cycle is quite adequate, especially because of the presence of vector instructions. Several other issues, in addition to those mentioned above, are also discussed and explored in the dissertation.