Global checkpointing for a concurrent processing system
RAJ SEKHAR PAMULA, Suchai Thanawastien, Y. L. VAROL · International Journal of Systems Science · 1990
In global checkpointing, N (N 1) checkpoints are maintained since Faults might not be detected on occurrence owing to the error latency time. The time required for rollback and recovery is therefore dependent on error latency, the number of checkpoints maintained, the intercheckpoint interval and the way checkpoints are selected for rollback. This paper presents a model for describing the error latency time distribution, which is pivotal in approximately determining the point of occurrence of a fault, given that it is delected. Two rollback schemes, sequential and lookahead, are investigated. Although the sequential rollback is easier to implement, the lookahead rollback has a smaller error recovery cost. Numerical results are also given for a system with two processes. The optimal number of checkpoints to be maintained by the system to minimize the cost of rollback for each of the rollback strategies is given.