A modular design approach to software fault-tolerance in distributed computer systems

E.-K. Park · 1988

In this dissertation, an approach to design fault-tolerant software which must possess behavior that is very reliable (complying with its design specifications) in distributed computer systems (DCS) is proposed. The DCS is modeled as a set of communicating sequential processes with constraints on their execution time. Each process corresponds to the execution of software components which are part of the distributed software system. The approach is based on decentralized protocol to monitor the behavior of software components distributed over processors and communicating among them. In this approach, the Distributed Fault Handler (DFH) which is distributed over processors is developed for error detection and recovery during the execution of DCS Software. An error classification scheme is discussed and an error detection technique is presented. An asynchronous rollback recovery is proposed which minimizes the rollback distance. This approach is useful for constructing general distributed applications software with fault tolerance capability, especially for applications which require high reliability. The asynchronous schemes proposed in the literature either are incomplete in that they lack recognition of unnecessary recovery points, or do not consider communication delay well, or restrict the application domain to transaction systems. The notable features of our approach include the domino-effect-free asynchronous scheme, and the dependency relationship information used to keep track of the consistent global state of the system; thus, we can find the consistent Recovery-Line quickly to provide a fast recovery. Also, the overhead involved in saving checkpoints and data or messages is comparatively low in our approach. The synchronization overhead in our approach is almost nonexistent compared with other approaches. Implementing our recovery block requires suitable entrance and acceptance block tests which can be obtained easily from system specifications. Our approach is relatively easy and inexpensive to implement, since the algorithm requires only the manipulation of data structures, Module Supervisor, Process Supervisor, and checkpoints in DFH with the recovery-block based error detection.

Read the paper · More papers on PaperTik