TREFU: An Online Error Detecting and Correcting Fault Tolerant GPGPU Architecture
K K Raghunandana, Varaprasad B.K.S.V.L., Matteo Sonza Reorda, Virendra Pal Singh · 2023
General Purpose Graphics Processing Units (GPGPUs) are extensively used in high-performance applications/systems, whose execution times may vary from a few days to months. Many times, these systems are expected to provide high reliability and availability. On the other hand, the high-throughput GPGPUs are fabricated with the latest cutting-edge technology. The shrinking transistor feature size and aggressive voltage scaling resulted in increased susceptibility to soft errors. Hence, GPGPU execution results cannot be trusted. This necessitates the employment of error detection and correction methods for reliable results. To mitigate soft error effects in the GPGPU execution pipeline, we propose a fault-tolerant microarchitecture called Triple modular Redundant Execution with idle Functional Units (TREFU) to detect and correct errors online. The proposed method is transparent to the application software. A new microarchitecture structure, replay buffer, is introduced to store temporary operands and results and used as a checkpoint. On error detection, the data in the duplicate copy of the replay buffers are used for Triple Modular Redundant (TMR) execution and error correction. The effectiveness of TREFU is demonstrated through the ISPASS 2009 and RODINIA benchmarks. TREFU's performance and power overheads are evaluated for an error-free run and at various error rates of executed instructions ranging from 1 to 50K. The simulation results show that complete error detection and correction across all threads can be achieved with a mean performance overhead of 4%, an average power overhead of 4%, and a peak power overhead of 5%.