A blocked implementation of Level 3 BLAS for RISC Processors

MJ Daydé, IS Duff · 1996

We describe a version of the Level 3 BLAS which is designed to be efficient on RISC processors. All our codes are written in Fortran and use loop-unrolling, blocking, and copying to improve the performance. A blocking technique is used to express the BLAS in terms of operations involving triangular blocks and calls to the matrix-matrix multiplication kernel (GEMM). This blocked implementation uses the same blocking ideas as in our earlier work on parallel implementations of the BLAS, although here the ordering of loops is designed for efficient reuse of data held in cache and not necessarily for parallelization. A parameter which controls the blocking allows efficient exploitation of the memory hierarchy on the various target computers. We present results on a range of RISC-based workstations and multiprocessors : 1. CRAY T3D 2. DEC 255 4/233 and DEC 8400 5/300 3. HP 715/64 4. IBM RS/6000-750 and IBM SP2 5. MEIKO CS2-HA 6. SGI Power Challenge 10000 7. SUN UltraSPARC-1 ...

Read the paper · More papers on PaperTik