Achieving Portable High Performance for Iterative Solvers on Accelerators
Karl Rupp, Philippe Tillet, Ansgar Jüngel, Tibor Grasser · PAMM · 2014
Abstract We propose performance enhancements for the implementation of the conjugate gradient method and the generalized minimum residual method for accelerators such as graphics processing units. Through a minimization of memory transfers from global memory via pipelining as well as a reduction of the number of compute kernels through kernel fusion, the performance is improved by up to two‐fold when compared to standard implementations based on vendor‐tuned routines. (© 2014 Wiley‐VCH Verlag GmbH & Co. KGaA, Weinheim)