Extended recovery protocol in distributed systems
H. Higaki, Makoto Takizawa · 2002
This paper proposes a novel protocol for taking checkpoints and asynchronously restarting the processes for the recovery from the transient faults in asynchronous distributed systems. In the protocol, each process can be restarted asynchronously without the livelock. Each process can have multiple checkpoints to minimize the amount of computation wasted by the recovery. Moreover the garbage collection method is discussed. Each process has at most n checkpoints where n is the number of the processes. Only O(l) control messages are required to be transmitted where l is the number of communication channels in the system.