On reconfigurability of some regular architectures.
Amiyaranjan Nayak · 1991
Fault tolerance is the survival attribute of computer systems; when a system is able to recover automatically from internal faults without suffering an externally perceivable failure, the system is said to be fault-tolerant. A common approach for achieving fault tolerance is through the incorporation of redundancy. The effectiveness of a redundancy-based fault tolerance scheme cannot simply be measured by the amount of redundancy built into the system, but also by the ability of the system to reconfigure in the presence of faults and to carry on its operation. Therefore, redundancy and reconfigurability are the two key aspects of fault tolerance that are directly related to reliability and low yield problems. Although many researchers have come forward to present their views in solving problems of low yield, the importance of the spare resources and fault distribution and their impact on reconfigurability have not been thoroughly investigated. This thesis studies the characteristics of catastrophic fault patterns; that is, patterns of faults whose occurrence can have catastrophic effects on the system and cannot be overcome regardless of how intelligent is the employed reconfiguration algorithm. Limits are established on the reconfigurability of regular networks such as one-dimensional systolic arrays, two-dimensional systolic arrays, and ring topologies.