Parallel performance of higher-order methods on GPU hardware
Tyler Spilhaus, Jared Buckley, Gaurav Khanna · IEEE International Conference on High Performance Computing, Data, and Analytics · 2015
There is considerable current interest in higher-order methods and also large-scale parallel computing in nearly all areas of science and engineering. In this work, we take a number of basic finite-difference stencils that compute a numerical derivative to different orders of accuracy and carefully study the overall performance of each, on a many-core processor i.e. a graphics processing unit (GPU). We conclude that if one has a code that exhibits a high order of convergence, then there is likely to be only a modest gain through GPU parallelism in the context of total execution or wall-clock time. Conversely, for a low order code that exhibits good parallel performance, there is insignificant gain through the implementation of a higher-order convergent algorithm.