The impact of HPF data layout on the design of efficient and maintainable parallel linear algebra libraries

IL (United States) Argonne National Lab., C Bischof, USDOE, Washington, DC (United States), Steven Huss‐Lederman, Elaine M. Jacobson, Xiaobai Sun, Anna Tsao · 1994

In this document, the authors are concerned with the effects of data layouts for nonsquare processor meshes on the implementation of common dense linear algebra kernels such as matrix-matrix multiplication, LU factorizations, or eigenvalue solvers. In particular, they address ease of programming and tunability of the resulting software. They introduce a generalization of the torus wrap data layout that results in a decoupling of {open_quotes}local{close_quotes} and {open_quotes}global{close_quotes} data layout view. As a result, it allows for intuitive programming of linear algebra algorithms and for tuning of the algorithm for a particular mesh aspect ratio or machine characteristics. This layout is as simple as the proposed HPF layout but, in the authors opinion, enhances ease of programming as well as case of performance tuning. They emphasize that they do not advocate that all users need be concerned with these issues. They do, however, believe, that for the foreseeable future {open_quotes}assembler coding{close_quotes} (as message-passing code is likely to be viewed from a HPF programmers` perspective) will be needed to deliver high performance for computationally intensive kernels. As a result, they believe that the adoption of this approach not only would accelerate the generation of efficient linear algebra software libraries but also would accelerate the adoption of HPF as a result. They point out, however, that the adoption of this new layout would necessitate that an HPF compiler ensure that data objects are operated on in a consistent fashion across subroutine and function calls.

Read the paper · More papers on PaperTik