Cooperative checkpointing theory

Adam J. Oliner, Larry Rudolph, Ramendra K. Sahoo · 2006

Cooperative checkpointing uses global knowledge of the state and health of the machine to improve perfor-mance and reliability by dynamically deciding when to skip checkpoint requests made by applications. Using results from cooperative checkpointing theory, this pa-per proves that periodic checkpointing is not expected to be competitive with the offline optimal. By leverag-ing probabilistic information about the future, coopera-tive checkpointing gives flexible algorithms that are op-timally competitive. The results prove that simulating periodic checkpointing, by performing only every dth checkpoint, is not competitive with the offline optimal in the worst case; a simple modification gives a prov-ably competitive algorithm. Calculations using failure traces from a prototype of IBM’s Blue Gene/L show an application using cooperative checkpointing may make progress 4 times faster than one using periodic check-pointing, under realistic conditions. We contribute an approach to providing large-scale system reliability through cooperative checkpointing and techniques for analyzing the approach. 1.

Read the paper · More papers on PaperTik