Transparently Resilient Task Parallelism for Chapel

Konstantina Panagiotopoulou, Hans‐Wolfgang Loidl · 2016

Hardware failure in High-Performance Computing systems is the norm. Failure data, collected over a nine year period across 22 large-scale systems of up to few thousands of NUMA or SMP nodes, at Los Alamos National laboratories, show averages of 20-1000 failures per year. This paper describes the design and outlines the implementation of transparent resilience for task parallelism in Chapel, a high-performance language developed for productive parallel programming. We detail the design directions and we implement a transparent resilience mechanism within Chapel's runtime system. Our primary goal is to ensure program termination in the presence of hardware failure of one or multiple nodes in the system. We evaluate our implementation using a set of five synthetic microbenchmarks covering Chapel's task parallel constructs and we quantify and discuss the small overheads and speedups noted for the resilient implementation compared to the latest non-resilient Chapel release.

Read the paper · More papers on PaperTik