Fault tolerance in distributed computing systems and databases

Seyed Hossein Hosseini · 1982

The problem of achieving fault-tolerance in a network of interconnected processing elements in which there is no central facility available for control, coordination, or mediation among the processing elements is considered. Past work has, in general, viewed fault tolerance as consisting of the four following steps, (i) error detection, (ii) fault location, (iii) system reconfiguration, and (iv) recovery of the system state from a faulty state to a fault-free state. The work in this thesis employs an earlier proposed model, in which it is assumed that some of the processing elements can test certain of the other processing elements with which they share a direct communication path for the presence of failures. Based on this model, distributed fault-diagnosis algorithms are developed such that every fault-free processing element can eventually, independently, and correctly diagnose the condition (faulty or fault-free) of all other network facilities. These algorithms represent a considerable extension to previous work in that they operate in the context where processing elements and communication paths between them may not only fail but may also be repaired, and returned to service, or replaced by fault-free facilities. It is also allowed that entirely new facilities may be added to the network. Using the same basic network model, a model of distributed error recovery strategy is proposed. This strategy integrates the distributed fault-diagnosis algorithm developed in the thesis, along with a technique for logical isolation of failed facilities and backward error recovery (roll back and retry), into an overall algorithm for insuring the correct performance of the computations in the modeled system. Finally, the problem of application of the hardware fault-tolerance in distributed database systems is considered as an example of the application of the results of the thesis. Techniques are given via which a distributed database can tolerate failures of some of its constituent facilities while maintaining the consistency and integrity of transactions carried out on the resources of the data base.

Read the paper · More papers on PaperTik