A methodology for the design and analysis of fault-tolerant operating systems
Marius Doru Soneriu · 1981
A methodology for the design and analysis of operating systems which can perform their specified functions in the presence of hardware and softward faults, to be called here Fault-Tolerant Operating Systems (FTOS), is introduced in this dissertation. The methodology covers the following main topics: modeling, time overhead analysis, comparative study of techniques for fault-tolerance, and reliability estimation. A general model for FTOS defined as a hierarchy of nested fault-tolerant machines is first presented. The model is obtained starting from a hierarchically structured conventional operating system. In order to transform it into a fault-tolerant system, each conventional machine is augmented with an Error Detection and Recovery (EDR) mechanism, thus obtaining a corresponding fault-tolerant machine. The EDR mechanism makes a conventional machine fault-tolerant by transforming its conventional operations into fault-tolerant operations. A model for fault-tolerant operations is thus developed, which can represent most techniques for fault-tolerance. Next, the applicability of the model for fault-tolerant operations is demonstrated by using it to investigate methods for reducing the time overhead introduced by redundancy. A new technique for fault-tolerance with low time overhead, called tandem, is presented. The technique is general, i.e. for certain values of a parameter N it degenerates into known techniques for fault-tolerance. A taxonomy for techniques for fault-tolerance in operating systems is then established. In order to define the taxonomy, a set of attributes which can fully characterize each technique for fault-tolerance and establish its placement within the taxonomy is identified. The distribution of techniques within the taxonomy is analysed; their proximity is related to similar features and a pattern of basic concepts is identified. Finally, a method for estimating the reliability of FTOS is introduced. By partitioning the FTOS into several fault-tolerant machines, the reliability of the entire system can be estimated as a function of the reliabilities of the individual machines. It is shown how the reliability of a general fault-tolerant machine can be estimated using a discrete-state, continuous-time Markov model.