An integrated study of fault tolerance in computing systems.

Tein-Hsiang Lin, Kang Geun Shin · Deep Blue (University of Michigan) · 1988

A general framework for the design and analysis of distributed fault-tolerant systems is proposed including fault/error occurrence and detection, error propagation, fault location, retry, system reconfiguration, damage assessment, and error recovery. Detection mechanisms are usually assumed to be so perfect that problems within a particular phase of fault-tolerance can be studied without considering its interplay with other phases. This dissertation shows that the assumption of imperfect detection mechanisms will greatly influence fault diagnosis, rollback recovery, and checkpointing. Under the imperfect detection assumption, error propagation becomes a critical problem in almost all phases of fault-tolerance. Therefore, we first develop a stochastic model for error propagation with a digraph to represent a distributed system and error propagation times between adjacent nodes as the basic model parameters. Algorithms are developed to calculate the error propagation times between any pair of nodes systematically and efficiently. Experiments are also performed to measure the basic parameters on an existing system. The model is then extended to describe the detection of faults and errors. Based on this model, we formulate and solve the problem of locating a faulty node, determining the optimal diagnostic order and the optimal diagnostic time for each node of the system upon the detection of an error. This model is also used for damage assessment, or estimation of the times of error propagation into individual nodes. Using the result of damage assessment, a problem of rollback recovery with checkpointing of concurrent processes is formulated and solved, determining the optimal rollback distance for each process. Two additional related problems are studied in the dissertation. One is concerned with the use of retry following a fault detection and the other with the optimal placement of checkpoints in a real-time task with or without the perfect detection assumption. A fault classification scheme is developed for on-line estimation of fault parameters. The proposed retry policy is based on Bayesian decision theory which does not require fault parameters to be known a priori and adapts itself to the changes in parameters during normal operation.

Read the paper · More papers on PaperTik