Efficient Intra-node Hierarchical Parallelisms And Dynamic Load Balancing Strategies On Heterogeneous Systems
Dikshant Pratap Singh, Mathialakan Thavappiragasam, Brice Videau · 2025
The way computing nodes are utilized can significantly impact the performance of applications on heterogeneous systems and single multi-GPU systems, such as the enterprise AI-focused DGX platform, which is currently trending for accelerated computing. The heterogeneity of the computing node/platform requires efficient intra-node parallelism to obtain high performance and energy efficiency. Additionally, varying heterogeneity over different HPC systems requires performance portability of the software tools. Hence, in this study, we are motivated to develop a) efficient dynamic load balancing multi-gpu scheduling strategies and b) node-to-thread level hierarchical parallelism methodologies using performance portability programming models, specifically OpenMP. We introduce and evaluate strategies for this hybrid parallelism and load balancing within a node/platform using OpenMP, MPI, and CUDA. For our evaluation, we have chosen Argonne Leadership Computing Facility’s (ALCF) NVIDIA-A100-based systems Polaris and DGX, and two applications: compute-and data-intensive. Our studies have shown that a) adaptive scheduling with multi-level task refinement effectively ensures load balancing across multiple GPUs. b) OpenMP+CUDA yields better performance, achieving a speedup between 1.9x and 2.1x, and OpenMP+OpenMP Offload delivers competitive performance, with a speedup ranging from 1.72x to 2.35x when compared to MPI+OpenMP Offload for the compute-intensive application. We implemented several optimizations that improved the performance by ≈ 10x. We also discuss the challenges encountered and potential solutions for intra-node multi-GPU utilization.