A checkpoint protocol for an entry consistent shared memory system

Nuno Neves, Miguel Castro, Paulo Cezar Camargo Guedes · 1994

Workstation clusters are becoming an interesting alter-native to dedicated multiprocessors. In this environment, the probability of a failure, during an application’s exeeution, increases with the execution time and the number of work-stations used. If no provision is made for handling failures, it is unlikely that long running applications will terminate successfully. One solution to this problem is process check-pointing. This paper presents a checkpoint protocol for a multi-threaded distributed shared memory system based on the en-try consistency memory model. The protocol allows trans-parent recovery from single node failures and, in some cases, from multiple node failures. A simple mechanism is used to determine if the system can be brought to a consistent state in the event of multiple machine crashes. The protocol keeps a distributed log of shared data ac-cessesin the volatile memory of the processes, taking advan-tage of the independent failure characteristics of workstation clusters. Periodically, or whenever the log reaches a high-water mark, each process checkpoints its state, independently from the others. The protocol needs no extra messagesdur-ing the failure-fke period, since atl checkpoint control in-formation is piggybacked on the memory coherence protocol messages. 1

Read the paper · More papers on PaperTik