An experimental and analytical validation of two fault-tolerant techniques in a tightly coupled network
Shin Heu · 1986
Since the recovery block (RB) was first introduced as a well-structured mechanism for facilitating software fault tolerance in basic sequential programming environments, several extensions of the RB have been proposed for use in distributed computing environments. To date, the cost effectiveness of such schemes has not been fully analyzed. A promising extension of the RB devised for real-time applications that need a recovery scheme which is faster than the conventional backward error recovery scheme is the scheme called the distributed recovery block (DRB). The other promising extension of the RB devised to facilitate process cooperation for recovery in distributed computing environments is the scheme called the conversation. These two extensions are complementary. The focus of the research reported here has been to examine the execution overhead of the two promising RB extensions incorporated into real-time tightly coupled networks (TCNs). The DRB scheme was originally proposed by Kim as a technique for unified treatment of hardware and software faults in real-time computing systems and for achievement of fast recovery by exploiting concurrent execution of alternate try blocks. Its implementation problems were not definitely solved; therefore, the research reported here has explored a strategy for optimal implementation of the DRB and employed experimental implementation as a way to validate the logical soundness and to evaluate the performance of the strategy formulated. A 6-node TCN of macro-data-flow structure was used as the hardware facility in the experimental implementation and the measured data are indicative of the high performance of the DRB scheme. The conventional scheme was proposed in abstract form by Randell as an approach to structuring fault-tolerant cooperating processes. Although its mechanization and implementation schemes have been studied to some extent, its potential advantages and disadvantages have not been fully explored. The use of the conversation increases execution time costs due to the tight synchronization imposed among participant processes, the recovery, etc. Analytic evaluation of these time costs has been performed by the use of a queueing network model that can cover various application environments.