Timing-Fault-Tolerant Out-of-Order Processor
Masahiro Goshima, Naruki Kurata, Ryota Shioya, Shuichi Sakai · 2013
A decrease of the feature size of LSIs leads to an increase of the effect of random variation. Tech- niques that detect and recover from timing faults can solve this problem. Existing techniques, however, can only be applied to simplest scalar processors. This paper proposes a new technique that can be applied to complex out-of-order superscalar processors. The point is how to handle the faults that occurs inside of the Reorder Buffer (ROB) and the Load/Store Queue (LSQ). We examine these commitment modules in detail, introduce the notion of Point-of-No-Return (PNR), and make a more general design rule that says not start of the commitment but passing through PNR is the actual commitment of an instruction. Then, we separate the store buffer from inside of the LSQ to below the PNR. And, the effect of the faults is removed by initializing the pipeline above the PNR. This scheme enables handling of faults that occurs any part above the PNR including the ROB and the LSQ. The latency of the commitment is prolonged by the separate store buffer. However, the simulation results show that IPC degradation caused by it is no more than 0.7% on average.