Optimizing matrix operations on a parallel multiprocessor with a memory hierarchy

William Jalby, Ulrike Meier · OSTI OAI (U.S. Department of Energy Office of Scientific and Technical Information) · 1986

Memory organizations of supercomputers (CRAY 2, CEDAR) tend to become more and more complex, and correspondingly data management in these memories becomes a crucial factor for achieving high performance. We study here an architecture combining vector and parallel capabilities on a two-level shared memory structure. For this class of architecture, we analyze and optimize matrix multiplication algorithms so as to obtain high efficiency kernels which can be used for many numerical algorithms such as LU and Cholesky factorizations, as well as Gram-Schmidt and Householder orthogonal factorization schemes. The performance of such kernels on the Alliant FX/8 multiprocessor, as well as their application to the Gram-Schmidt procedure is described in detail.

Read the paper · More papers on PaperTik