Making TifaMMy fit for tomorrow: Towards future shared memory systems and beyond

Alexander Heinecke, Carsten Trinitis · 2011

In this paper, we present the recent port to and latest results of our cache-oblivious algorithms and implementations of parallel LU decomposition code TifaMMy on two new architectures: SGI's UltraViolet distributed shared memory machine, and Intel's latest x86 architecture Sandy Bridge. TifaMMy's matrix multiplication and LU decomposition routines have been further optimized with regard to these new architectures. Results are discussed and compared with Intel's architecture specific and optimized numerical Math Kernel Library (MKL) for both the standard C++ version with vectorization compiler switches and TifaMMy's highly optimized vector intrinsics version.

Read the paper · More papers on PaperTik