Implementing Dense Linear Algebra Algorithms Using Multitasking on the CRAY X-MP-4 (or Approaching the Gigaflop)
Jack J. Dongarra, Tom Hewitt · SIAM Journal on Scientific and Statistical Computing · 1986
This note describes some experiments on simple, dense linear algebra algorithms. These experiments show that the CRAY X-MP is capable of small-grain multitasking arising from standard implementations of $LU$ and Cholesky decomposition. The implementation described here provides the “fastest” execution rate for $LU$ decomposition, 718 MFLOPS for a matrix of order 1000.