CHECKPOINTING ALGORITHMS IN DISTRIBUTED SYSTEMS
Wei Xiao · Chinese Journal of Computers · 1998
Checkpointing can save and restore programs running state. It is thebackbone of certain program control utilities, such as process migration, fault-tol-erance, and playback debugging, etc. In the paper, existing checkpointing algo-rithms are classified and discussed in detail. Existing checkpointing algorithms areclassified into two broad categories. One is for single-process programs, the otheris for distributed programs. Moreover, the checkpointing algorithms for distributedprograms are classified into also-asynchronous checkpointing algorithms and consis-tent checkpointing algorithms. Furthermore, the typical methods to improve theperformance of checkpointing algorithms are also introduced in the paper, such asincremental checkpoints, CAME, compression, user-directed checkpoints, mainmemory, copy-on-write, and CLL, etc. The first four optimizations reduce theamount of information saved in a checkpoint, and the rest increase the concurrencyof checkpointing by overlapping executing a target program with writing check-points to the disk. Overhead and latency are used to evaluate the performance ofcheckpointing algorithms in the paper. In the end, the paper discusses the commondrawbacks of existing checkpointing algorithms and the future work.