An Analysis of Multilevel Checkpoint Performance Models
Daniel Dauwe, Sudeep Pasricha, Anthony A. Maciejewski, Howard Jay Siegel · 2018
Periodic application checkpointing to a parallel file system has for years been the standard strategy for providing high performance computing (HPC) systems with resilience to system failures. The traditional checkpoint/restart protocol has until recently proved sufficient for mitigating the impact of these failures on application performance. However, as system sizes approach exascale levels, the frequency of failures and the time required to checkpoint/restart an exascale-size application increases and the efficiency of traditional checkpointing decreases substantially, making it no longer a viable option for providing resilience to future HPC systems. The most frequently proposed solution for providing future systems with resilience has been to design multilevel checkpointing protocols. However, the relationship between system failure rates, checkpoint/restart overhead, and duration of time between successive checkpoints is complex and finding the optimal duration of time between successive checkpoints is an open and challenging problem. This work presents a novel execution time prediction model we have developed that takes into consideration execution events that have not been considered by previous multilevel checkpointing models. We show how this model can be used to select checkpoint intervals and demonstrate why consideration of these execution events is important. We validate our work through simulation and provide a comparison to several optimization strategies proposed in other work to demonstrate the advantage gained by considering these execution events.