Communication-Computation Overlapping with Dynamic Loop Scheduling for Preconditioned Parallel Iterative Solvers on Multicore and Manycore Clusters

Kengo Nakajima, Toshihiro Hanawa · 2017

Preconditioned parallel solvers based on the Krylov iterative method are widely used in scientific and engineering applications. Communication overhead is a critical issue when executing these solvers on large-scale massively parallel supercomputers. In this work, we introduced communication-computation (CC) overlapping with dynamic loop scheduling of OpenMP to the sparse matrix-vector multiplication (SpMV) process of a parallel iterative solver. We then used the solver to evaluate the performance of a parallel finite element application (GeoFEM/Cube) on multicore and manycore clusters. The dynamic loop scheduling of OpenMP improved the efficiency of CC overlapping in halo exchanges, and the developed method attained a significant performance improvement of 40-50% for parallel iterative solvers in strong scaling using up to 16,384 cores of a Fujitsu PRIMEHPC FX10 supercomputer and an Intel Xeon Phi (KNL) cluster. Finally, the developed method was applied to GeoFEM/Cube using a parallel BiCGSTAB solver with sparse approximate inverse (SAI) preconditioning, and a 15-20% performance improvement was obtained on 12,288 cores of the Fujitsu FX10 and the KNL cluster.

Read the paper · More papers on PaperTik