Optimized Execution of Parallel Loops via User-Defined Scheduling Policies
Seonmyeong Bak, Yanfei Guo, Pavan Balaji, Vivek Sarkar · 2019
On-node parallelism continues to increase in importance for high-performance computing and most newly deployed supercomputers have tens of processor cores per node. These higher levels of on-node parallelism exacerbate the impact of load imbalance and locality in parallel computations, and current programming systems notably lack features to enable efficient use of these large numbers of cores or require users to modify codes significantly. Our work is motivated by the need to address application-specific load balance and locality requirements with minimal changes to application codes.