Heterogeneous Parallel Computing Platforms and Tools for Compute‐Intensive Algorithms: A Case Study
Daniele D’Agostino, Andrea Clematis, Emanuele Danovaro · 2014
This chapter analyzes some of the programming models and tools, both commercial and freely available, in terms of the provided support and achievable performance with respect to some widely used compute-intensive algorithms such as the N-body and the convolution algorithms. It focuses on providing a clear measure of the different efficiency figures with respect to the programming paradigm considered, the adopted tool, and the achievable peak performance. It is quite easy to achieve an order of magnitude speed improvement for selected compute kernels with a limited effort, but also that developers have to know the specific features of every compute capability to achieve much higher performance. The OpenACC application program interface describes a collection of compiler directives to specify loops and regions of code in standard C, C++, and Fortran to be offloaded from a host CPU to an attached accelerator, providing portability across operating systems, host CPUs, and accelerators.