Reconfigurable NoC and Processors Tolerant to Permanent Faults

Alirad Malek · Chalmers Research (Chalmers University of Technology) · 2015

Advances in semiconductor industry have led to reduced transistor dimensions andincreased device density, but inevitably they have compromised the reliability of moderncomputing systems. In this thesis, we address the reliability problemby exploitinghardware reconfiguration for tolerating permanent faults. Processing components in asystem-on-chip are divided into smaller Substitutable Units (SUs) and reconfigurableinterconnects are used to isolate defective SUs and connect spare units to create afault-free component. Furthermore, employing fine-grain logic for instantiating a functionallyequivalent unit is another reconfiguration option considered. Based on theseapproaches, the first part of this thesis presents a probabilistic analysis of reconfigurabledesigns for calculating the average number of constructable components at differentfault densities. Considering the area overheads of reconfigurability, we evaluate theresilience of various reconfigurable designs with different granularities (SU sizes). Concisely,the results reveal that the combination of fine and coarse-grain reconfigurationoffers up to 3£ more fault-tolerance compared to component redundancy. Performinga design-space exploration to find the most efficient granularity mix shows thatdifferent fault densities require different granularities of substitutable units to maximizefault-tolerance. Moreover, we explored the performance effects of pipelining thereconfigurable interconnects in adaptive processors and observed that the operatingfrequency and execution time of pipelined design is roughly 2.5£ and 2£ better than thedesign with non-pipelined interconnects, respectively. In the second part of this thesis,we describe RQNoC, a service-oriented Network-on-Chip (NoC) resilient to permanentfaults. We characterize the network resources based on the particular service they supportand, when faulty, bypass them allowing the respective traffic class to be redirected.We propose service merging (SMerge) and service detouring (SDetour) as the two serviceredirection schemes. Different RQNoC configurations are implemented and evaluatedin terms of performance, area, power consumption and fault tolerance. Concisely, theevaluation results show that compared to the baseline network, SMerge requires 51%more area and 27% more power and has a 9% slower clock but maintains at least 90% ofthe network connectivity even in presence of 32 permanent network faults.

Read the paper · More papers on PaperTik