FAULT TOLERANCE IN PACKET COMMUNICATION COMPUTER ARCHITECTURES

Chun-Ming Leung · 1980

It is attractive to implement a large scale parallel processing system as a self-timed hardware system with decentralized control and to improve maintainability and availability in such a system through fault tolerance. In this thesis we show how to tolerate hardware failures in a self-timed hardware system with a packet communication architecture, designed to execute parallel programs organized by data flow concepts. We first formulate a design methodology for incorporating redundant hardware into self-timed systems for fault tolerance. Redundancy management problems in self-timed systems are illustrated with a byte-sliced hardware module structure. Robust algorithms are given for synchronizing byte slices in a redundant module so that their outputs can be decoded to detect and/or mask hardware failures. Hardware implementation of these redundancy management algorithms is studied under a stuck-at fault model, a random pulse train fault model and a random wave train fault model. In studying the design of fault-tolerant data flow processors we have also developed a dynamic redundancy scheme for masking hardware failures in a multiprocessor architecture designed to execute parallel programs organized by data flow principles. Novel features of this architecture include use of packet networks to support communication among processing elements and dynamic allocation of a homogeneous set of functional units to service requests. Program organization and hardware module designs to support the dynamic redundancy scheme are described.

Read the paper · More papers on PaperTik