On the reliability of an object based distributed system (resource sharing)

Fawzi F. Gherfal · 1985

One of the main advantages of distributed computer systems is the ability to share resources among heterogeneous autonomous single systems. The distributed system can be viewed as a number of resources provided by the collection of single systems. A number of resources may be combined to construct a distributed resource or may cooperate to implement a distributed application or subsystem. A number of failures, such as a node crash, link failure, and network partition, in addition to concurrent access of the resource, may cause system state inconsistency and resource unavailability, rendering the distributed system unreliable. In this dissertation we design an efficient software support for reliable and modular resource sharing in this distributed environment. An object based reliability model is proposed that captures the properties of the environment of interest and provides for reliable resource sharing in spite of failure. The model is based on an abstract object called the Recoverable Module that models a resource which is tolerant to failure and able to manipulate exceptions. The Recoverable Module forms the basic construct for building the reliable distributed system. In addition, the Recoverable Module forms the unit of recovery and controls synchronization. We believe that such an approach is more superior in terms of containing failure effects and achieving more efficient recovery than the approaches that are proposed in transaction based recovery models, where the entire distributed transaction or subtransaction is used as the unit of synchronization and recovery. One important feature of the model is that it allows for the specification of resource dependent recovery and synchronization semantic knowledge. Such knowledge aids in achieving flexibility, a high degree of concurrency, and efficient recovery. A number of mechanisms are proposed to implement the reliability properties of the model. These mechanisms include: (1) Node recovery, to detect, isolate, and provide for the recovery of a crashed node. (2) Module Recovery, to restore the Module to a consistent state and localize the effects of failure. (3) Local Module synchronization mechanism, to express and reserve the Module's dependent synchronization constraints and synchronization semantic knowledge. (4) An Optimistic concurrency mechanism, to synchronize multiple transactions and increase the degree of concurrency. (5) An exception handling mechanism that allows for specification of the Module's dependent recovery semantic knowledge and control error propagation.

Read the paper · More papers on PaperTik