ParaGraph: Weighted Graph Representation for Performance Optimization of HPC Kernels
Ali TehraniJamsaz, Alok Mishra, Akash Dutta, Abid Muslim Malik, Barbara Chapman, Ali Jannesari · 2024
GPU-based HPC clusters are attracting more sci-entific application developers due to their extensive parallelism and energy efficiency. In order to achieve portability among a variety of multi/many core architectures, a popular choice for an application developer is to utilize directive-based parallel programming models, such as OpenMP. However, even with OpenMP, the developer must choose from among many strategies for exploiting a GPU or a CPU. This paper introduces a new graph-based program representation for optimization of OpenMP applications. The originality of this work lies in the augmentations of Abstract Syntax Trees (ASTs) and the introduction of edge weights to account for loop and condition information. We evaluate our proposed representation by training a Graph Neural Network (GNN) to predict the runtime of OpenMP code regions across CPUs and GPUs. Various transformations utilizing collapse and data transfer between the CPU and GPU are used to construct the dataset. The trained model is used to determine which transformation provides the best performance. Results indicate that our approach is effective and has normalized RMSE as low as$4\times 10^{-3}$to at most$1\times 10^{-2}$in its runtime predictions.