Load-Balanced LU and QR Factor and Solve Routines for Scalable Processors with Scalable I/O
Jean-Philippe Brunet, Palle Pederson, S. Lennart Johnsson · Digital Access to Scholarship at Harvard (DASH) (Harvard University) · 1994
. The concept of block--cyclic order elimination can be applied to out--of-- core LU and QR matrix factorizations on distributed memory architectures equipped with a parallel I/O system. This elimination scheme provides load balanced computation in both the factor and solve phases and further optimizes the use of the network bandwidth to perform I/O operations. Stability of LU factorization is enforced by full column pivoting. Performance results are presented for the Connection Machine system CM--5. 1 Introduction Load--balance for in--core matrix factorization on distributed memory architectures can be achieved using a cyclic ordering of the data. In fact, one need not allocate data explicitly in a cyclic fashion. Instead, the elimination can be performed in a cyclic order. That block--cyclic order elimination is an efficient alternative to block--cyclic data allocation for in--core dense matrix factorization was shown in [1]. The present note extends this concept to out--of--core ...