Recovery in Multicomputers with Finite Error Detection Latency

Parimala Venkata Krishna, Nitin H. Vaidya, Dhiraj K. Pradhan · 1994

In most research on checkpointing and recovery, it has been assumed that the processor halts immediately in response to any internal failure (fail-stop model). This paper presents a recovery scheme (independent checkpointing and message logging) for a multicomputer system consisting of processors having a non-zero error detection latency. Our scheme tolerates bounded error detection latencies, thus, achieving a higher fault coverage. The simulation results show that for typical detection latency values, the recovery overhead is almost independent of the detection latency.

Read the paper · More papers on PaperTik