AN EFFICIENT COORDINATED CHECKPOINTING APPROACH FOR DISTRIBUTED COMPUTING SYSTEMS WITH RELIABLE CHANNELS

Lalit Kumar Awasthi, Manoj Kumar Misra, Ramesh Chander Joshi · International Journal of Computers and Applications · 2012

AbstractIn distributed systems, likelihood of failure increases with increase in the number of processes and a single failure often renders the entire system state useless. Checkpointing and rollback recovery is a common technique used for increasing the system reliability against various anticipated and unanticipated failures. Checkpointing can be independent, quasi-synchronous and coordinated. Coordinated checkpointing can be blocking or non-blocking. Also, either all the processes in the distributed system may need to checkpoint or only a minimum number of processes may be required to checkpoint. Minimizing the number of processes to checkpoint may introduce blocking. The non-blocking checkpointing protocols introduce overhead of piggybacking some information for non-intrusiveness. Minimization of this piggybacked information is the objective of our work. We have designed a non-blocking coordinated checkpointing protocol for distributed systems with reliable communication channels that minimize piggyba...

Read the paper · More papers on PaperTik