Extra High Speed Matrix Multiplication on the Cray-2

David A. Bailey · SIAM Journal on Scientific and Statistical Computing · 1988

The Cray-2 is capable of performing matrix multiplication at very high rates. Using library routines provided by Cray Research, Inc., performance rates of 300 to 425 MFLOPS can be obtained on a single processor, depending on system load. Considerably higher rates can be achieved with all four processors running simultaneously. This article describes how matrix multiplication can be performed even faster, at up to twice the above-listed rates. This can be achieved by: (1) employing Strassen’s matrix multiplication algorithm to reduce the number of floating-point operations performed and (2) utilizing local memory on the Cray-2 to avoid performance losses due to memory bank contention. The numerical stability and potential for parallel application of this procedure are also discussed.

Read the paper · More papers on PaperTik