Fault-tolerant real-time objects

K. H. Kim, C. Subbaraman · Communications of the ACM · 1997

Both the complexity of large-scale real-time (RT) computer systems in safety-critical application fields and the reliability expectations of the user community for such computer systems have been growing fast in recent years. Ideally, fault detection and recovery actions must always be executed such that intended output actions of real-time computations take place on time. Such an idealistic type of fault tolerance, which is to accomplish all critical actions successfully in spite of component failures, is called the action-level fault tolerance. Desirable fault tolerance techniques must be scalable in that they must be applicable to various distributed and/or parallel computer systems of different sizes. This article briefly reviews some recently established approaches for extending the conventional object structuring scheme into a powerful scheme for structuring both real-time and non-RT computations. The basic principles presented here may be largely applicable to environments where conventional object structuring approaches are used but validating this conjecture appears to be a meaningful topic for future research.

Read the paper · More papers on PaperTik