OpenMP’s Asynchronous Offloading for All-pairs Shortest Path Graph Algorithms on GPUs
Mathialakan Thavappiragasam, Vivek Kale · 2022
Numerical scientific computations, which are based on floating-point operations, have been sped up greatly via GPUs or other accelerators of supercomputers. However, combinatorial scientific computations, which are based on integer operations, do not use GPUs on a node well. The reason for this is that offloading of data and computation from the CPU (host) to a GPU (accelerator device) of a node of a supercomputer is by default synchronous. Synchronous offloading is costly if the host can do other meaningful computation or offload other independent tasks on the accelerator. To counter these costs, the capability of asynchronous offloading in GPU programming models is used by an application programmer to overlap data transfer and computation on the host when she deems it to be correct and performant. The OpenMP and CUDA programming models, and their respective vendor or platform HPC software, i.e., compiler suite, provide this capability to users in different ways. A challenge is in understanding the use of and assessing the effectiveness of this asynchronous offloading capability of OpenMP for scientific applications dominated by integer operations. This work shows and assesses the use of such a capability in the US Department of Energy LCF Floyd-Warshall benchmark implemented with OpenMP and with its vendor language counterparts (HIP and CUDA). By experimenting with our improved version of the benchmark using asynchronous OpenMP offloading, we found that using IBM’s xlc OpenMP on Summit was 1. 35x faster than using the NVIDIA’s nvcc CUDA, and using HPE’s cce OpenMP on Summit was l.llx faster than using the AMD’s hipcc HIP. These results provide insight on the use and the attainable performance of the asynchronous offload capabilities in OpenMP for such applications run on emerging supercomputers.