A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies

Jonas H. Müller Korndörfer, Ali Mohammed, Ahmed Eleliemy, Quentin Guilloteau, Reto Krummenacher, Florina M. Ciorba · IEEE Access · 2025

Scientific and data science applications are becoming more complex, with increasingly demanding computational and memory requirements during execution. Modern high performance computing (HPC) systems offer increased parallelism and heterogeneity across nodes, devices, and cores. Thus, effective scheduling and load balancing techniques are crucial to maximize applications performance on such systems. Commonly used node-level parallelization frameworks, such as OpenMP, employ an increasing number of advanced scheduling algorithms to support various applications and HPC platforms. For arbitrary application-system pairs, this results in an instance of thescheduling algorithm selection problem, which requires fast and accurate selection methods for the diversity of applications’ workloads and computing systems’ characteristics. In this work, we study the problem of learning to select scheduling algorithms in OpenMP. Specifically, we propose and use expert-based and reinforcement learning (RL) based selection approaches. We assess the effectiveness of each approach through an extensive performance analysis campaign with six applications and three systems, exposing capabilities and limits. We found that learning the highest-performing scheduling algorithm using RL-based methods is effective, but at a high exploration cost, and that the most relevant factor is the type of reward used by the RL methods. We also found that expert-based selection indeed leverages expert knowledge with fewer exploration needs, albeit at the risk of not selecting the highest-performing scheduling algorithm for a given application-system pair. We combine expert knowledge with RL-based approaches and show improved performance. Overall, this work shows that the selection of scheduling algorithms during execution is possible and beneficial for OpenMP applications. We anticipate that this work can be extended and combined with the selection of scheduling algorithms for MPI-based applications (provided that there is a portfolio of scheduling algorithms on distributed memory nodes), thereby optimizing scheduling decisions across parallelism levels.

Read the paper · More papers on PaperTik