Crash recovery with little overhead

Tony Tong-Ying Juang, Subbarayan Venkatesan · 2002

Recovering from processor failures in distributed systems is an important problem in the design and development of reliable systems. Two solutions to this problem which involve very little overhead are presented. Without appending any information to the messages of the application program, it is shown that it is possible to recover from failures using O( mod V mod mod E mod ) messages where mod V mod is the number of processors and mod E mod is the number of communication links in the system. The second algorithm can be used to recover from processor failures without forcing nonfaulty processors to roll back under certain conditions.>

Read the paper · More papers on PaperTik