Efficient implementation of the distributed recovery block (DRB) scheme in highly parallel computer systems

Byoung-Joon Min · 1991

The Distributed Recovery Block (DRB) scheme is essentially an active redundancy scheme where multiple processors concurrently execute multiple versions of a software component and the same acceptance test to provide for real-time forward recovery capabilities. The objective of this research was to study implementation techniques of the scheme in highly parallel computer systems and experimentally evaluate the scheme in real-time computer network testbeds. Past research, both analytic and experimental, focused mainly on the case where the DRB scheme was applied to only one computing station. In the recent experiment conducted as an early part of this dissertation research, the DRB scheme was incorporated into two adjacent computing stations in order to further study implementation techniques and to obtain more accurate performance data. A high fidelity simulation of a real-time application was developed to run on a testbed called the Crossbar Multi-microcomputer System (CMS), which consists of 7 Zilog Z8001-based single-board microcomputers and shared memories. Several fundamental implementation issues that arise when the PRB scheme is applied to multiple computing stations were identified and the performance of the DRB scheme in terms of overhead and recovery time was measured. Since the DRB scheme does not require special hardware mechanisms, it can be incorporated into more general parallel computer systems to achieve fault tolerance. Mapping of parallel tasks to a parallel computer system such that each task is duplicated into two nodes to form a DRB station is called the full DRB mapping. In the case of hypercubes, a procedure for converting a given optimal task mapping into an efficient full DRB mapping was developed by Kim and Kavianpour. Mapping techniques which are highly useful in obtaining efficient full DRB mappings of various task graphs to known parallel computer systems such as circular linked arrays have been developed. In addition, a new architecture called the BC hybrid structure has been formalized and the motivation here has been to create an architecture suitable for many important space and defense applications which require efficient facilities for both broadcasting data among multiple processors and exchanging data between neighbor processors. The DRB mapping into this architecture has also been considered in this dissertation research. Therefore, this dissertation provides several techniques that are fundamental in obtaining fault-tolerant real-time parallel computer networks based on the DRB scheme. It also provides an experimental demonstration of the real-time fault tolerance capability of such a network.

Read the paper · More papers on PaperTik