Construction of Performance Model of Tile CAQR and Performance Result of the Implementation
Masatoshi Takayanagi, Tomohiro Suzuki · 2017
Highly parallel computational resources can be exploited by asynchronously executing many fine-grained tasks. The tile algorithm for matrix decomposition can generate many fine-grained tasks, so is suitable for modern multicore/manycore architectures. However, the performance of this algorithm significantly depends on the tile size. We implement the tile algorithm in OpenMP/MPI hybrid fashion on a cluster system and construct a performance model that tunes the tile size by measuring the performance of simple computational kernels in our implementation. In this report, we test our communication-avoiding tile QR implementation for tall and skinny matrices on the K computer, and demonstrate the applicability of the performance model.