Portable and Vendor-Independent Low-Level Programming and Performance Benchmarking for Graphics Cards and Processors

David Rohr, V. Lindenstruth · 2017

GPUs have the potential to speed up programs significantly and are one opportunity to increase the scientific reach of compute-intense scientific applications. Several new programming models based on C and other languages have evolved to leverage the potential of such parallel architectures. Still, the development of individual source code versions using different languages and APIs deteriorates the maintainability of the code. It can also lead to slightly different outputs complicating the verification of the results. For comparing the compute performance, the different nature of the processors, different pricing, and different speed grades of hardware must be taken into account. In this paper, we summarize our experience from adapting a set of applications to GPUs. We present several cases how we implement generic code for multiple architectures and how we overcome the challenges that occurred. The presented applications encompass an algorithm to reconstruct the trajectories of particles for the ALICE High Level Trigger at the Large Hadron Collider at CERN; the Linpack benchmark used for ranking the performance of supercomputers and in particular its matrix multiplication substep; Reed-Solomon based failure erasure coding for redundant data storage; Lattice Quantum Chromo Dynamics computations; and an application for evaluating electron microscopy images.

Read the paper · More papers on PaperTik