Two phase algorithm to establish consistent checkpoints for recovery in multi process environment
Ashish Chauhan · IOSR Journal of Engineering · 2014
The problem of bringing a distributed multi-process system to a consistent state after transient failure is quite a difficult task.This paper addresses two components of this problem by describing a distributed algorithm to create consistent checkpoints, as well as rollback-recovery algorithm to recover the system to a consistent state.The motive of this algorithm is to make system more fault-tolerant.The likelihood of fault grows as systems are becoming more complex and applications are requiring more resources, including execution speed, storage capacity and many more.A checkpoint is a local state of a process saved on stable storage.In case of fault in distributed system, checkpoints enable the execution of the program to resumed from a previous consistent global state rather than resuming the execution from the beginning.When a process takes a checkpoint, minimal number of additional processes is forced to take checkpoints.Similarly when a process rollback and restarts from failure, a minimal number of additional processes are forced to rollback with it.This paper presents an efficient that algorithm works on Coordinator and Cohorts mechanism where the consistent set of checkpoints is established when it is directed by the coordinator.The checkpoints are established when all cohorts finishes the task assigned by the coordinator and there is no message in transit.