On Checkpoint Latency
Nitin H. Vaidya · 1995
Checkpointing and rollback is a technique for minimizing loss of computation in presence of failures. Two metrics can be used to characterize a checkpointing scheme: (i) checkpoint overhead (increase in the execution time of the application because of a checkpoint) , and (ii) checkpoint latency (duration of time required to save the checkpoint). For many checkpointing methods, checkpoint latency is larger than checkpoint overhead. This paper evaluates the expression for "average overhead" of the checkpointing scheme as a function of checkpoint latency and overhead. It is shown that the "average overhead" is much more sensitive to the changes in checkpoint overhead, as compared to checkpoint latency. Also, for equi-distant checkpoints, the optimal checkpoint interval is shown to be independent of the checkpoint latency. 1 Introduction Many applications (sequential and parallel) require large amount of time to complete. Such applications can encounter loss of a significant amount of co...