An approach to design of fault-tolerant real-time tightly coupled networks and its experimental validation
Jaechul Yoon, K. H. Kim · 1988
With rapid improvement of cost and reliability of computer modules, and with the advent of the distributed computer system (DCS) which offers new ways of achieving fault-tolerant operation, computers are increasingly used in critical real-time applications. Software fault tolerance has been realized as a necessary feature to improve system reliability as software becomes more and more complex. Still, current software fault tolerance technology has a lot of room for improvement in supporting the needs of such critical applications. The objective of this research was to develop software fault tolerance techniques applicable to the distributed real-time applications and experiment with the techniques on real-time DCS testbeds. Tightly coupled real-time distributed network testbeds and distributed real-time applications were established to support validation of the software fault tolerance techniques studied in this research through experimentation. Effectiveness of software fault tolerance can be confirmed by testbed-based approaches, and the experimental results give insight into behavior of the algorithms and allow the selection of appropriate system parameters for optimizing system performance. The distributed recovery block scheme with non-disruptive rejoin capabilities (DRB/NDR) and the variant structures of conversation were explored in this research. These schemes, which are extensions of the recovery block (RB), were devised for distributed real-time applications. Implementation strategies of the DRB/NDR were studied and the scheme was experimented with a testbed to validate logical soundness and evaluate performance of the implementation strategy formulated. Viability of the conversation scheme as an approach to achieving fault tolerance in a real-time distributed computing environment was investigated. Recovery mechanisms from a temporary blackout, which may occur due to high energy events or instantaneous power failure, were studied and validated by using testbed-based approaches. These software fault tolerance techniques were applied to different levels of application environments. The temporary blackout recovery mechanisms can be applied to all the nodes of a network system. The DRB scheme can be applied to each node of the system and several of those nodes may participate in a conversation. Therefore, software fault tolerance can be effectively achieved by applying these techniques to appropriate levels of a network system. (Abstract shortened with permission of author.)