A strategy for running large scale applications based on a model that optimizes the checkpoint interval for restart dumps
John T. Daly · 2004
We describe a method of approximating the optimum checkpoint restart strategy for minimizing application run time on a system exhibiting Poisson single component failures and discuss an effective strategy for applying these results on architectures where the application runtime is long compared to the mean time between failures. We begin by developing a model akin to Young's (1974, Comm. of the ACM, 17, 530-531). Then we apply this model to a simple checkpointing strategy and demonstrate its impact on the wall clock time for application completion.