OpenMP Offload Features and Strategies for High Performance across Architectures and Compilers
Arijit Bhattacharjee, Christopher S. Daley, Ali Jannesari · 2023
High performance accelerated computing has dawned in a new era of highly specialized code that depends on the target architecture. All the latest pre-exascale and exascale class supercomputers are accelerated systems. Efficient exploitation of the target architecture requires programming approaches that can take advantage of the characteristics of the hardware. OpenMP 4 introduced directives designed with the parallel architecture of GPUs in mind, but these directives are often different from those required for highest performance on CPUs. We demonstrate that by using newer OpenMP 5.0 offload features such as the loop directive and metadirective as well as profile guided strategies, we can achieve performance benefits with productivity gains. In this paper we experiment with HPGMG, a benchmark in the SPEChpc 2021 suite, to evaluate these newer directives versus OpenMP 4.5 directives, comparing performance, productivity and portability across architectural targets and production compilers on NERSC-9 Perlmutter. Our implementation gives up to a $2.74\times$ speedup compared to the baseline performance on GPUs and a $4.4\times$ speedup while running in host fallback mode. In addition, our implementation requires 66% fewer lines of code in the key kernels.