Time-Based Coordinated Checkpointing

Nuno Neves · 1998

Distributed systems are being used to support the execution of applications ranging from long-running scientific simulators to e-commerce on the Internet. In this type of environment, the failure of one of its components, either a computer or the network, may prevent other components from completing their tasks. Since the probability of failure increases with the number of computers and execution time, it is likely that these applications will be interrupted unless provision is made for failure handling. In this thesis we address the problem of fault recovery in distributed systems. The thesis describes two variations of a coordinated checkpoint protocol that uses time to re-move most causes of overhead, and to avoid all types of direct coordination. The time-based pro-tocol does not have to transmit extra messages, does not need to tag the application messages, and only accesses the stable storage when the checkpoints are saved. The thesis also describes a new coordinated checkpoint protocol that is well adapted to mobile environments. It uses time to indi-rectly coordinate the creation of new global states, and it saves two different types of checkpoints to adapt its behavior to the current network characteristics. Traditional techniques for fault diagnosis in distributed systems, either based on watch-dogs or polling, exchange performance with detection latency. The thesis introduces a complementary

Read the paper · More papers on PaperTik