Try, try till you succeed: Multiple checkpointing and rollback in distributed systems
Vinit A. Ogale · 2004
We present a multiple checkpointing and recovery protocol for fault tolerance in distributed systems. This is useful in scenarios where a fault trigger may result in observable undesirable eects after indeterminate time, i.e., when dormant faults are observed. The primary assumption is that the fault trigger occurs in rare circumstances and it is highly probable that the fault will not reoccur in another run. To ecien tly collect and store global snapshots we propose an online algorithm for slicing a computation w.r.t. conjunctive predicates. In this paper we deal with tolerating software faults in distributed systems. Software faults are extremely dicult, if not impossible to avoid. Hence a distributed system which can tolerate software faults with minimum runtime overhead, is highly desirable. In this paper we consider faults which can be circumvented by re-executing the program. We also assume that faults can be detected eventually. Some common examples of such faults are channel faults, resource availability and bad states in probabilistic algorithms. We allow the delayed detection of such faults, i.e., the fault may be detected long after it has actually occured. We call these type of faults dormant faults. One of the popular techniques for fault tolerance is checkpointing and rollback. Many approaches have been proposed for global checkpointing and rollback in distributed systems [3]. Elnozahy et. al. [2] give a comprehensive survey of