Data prefetching and multilevel blocking for linear algebra operations

Juan J. Navarro, Elena García-Diego, José R. Herrero · 1996

Much effort has been directed towards obtaining near peak performance for linear algebra operations on current high performance workstations. The large amounts of data accesses however, make performance highly dependent on the behavior of the memory hierarchy. Techniques such as Multilevel Blocking (Tiling), Data Precopying, Software Pipelining and Software Prefetching have been applied in order to improve performance. Nevertheless, to our knowledge, no other work has been done considering the relation between these techniques when applied together. In this paper we analyze the behavior of matrix multiplication algorithms for large matrices on a superscalar and superpipelined processor with a multilevel memory hierarchy when these techniques are applied together. We study and model the performance and limitations of different codes. We also compare two different approaches to data prefetching, binding versus non-binding, and find the latter remarkably more effective than the former due...

Read the paper · More papers on PaperTik