Accurate Analysis of Silent Data Corruptions in Programmable AI Accelerator Microarchitectures
Odysseas Chatzopoulos, Maria Trakosa, Dimitris Gizopoulos · 2025
Programmable AI accelerators become increasingly important to modern computing infrastructure, thus, their reliability is critical for the integrity of the produced results. Silent Data Corruptions (SDCs)-incorrect program outputs that occur without any warning or notification-have been reported by hyperscalers such as Meta, Google, and Alibaba, affecting both CPUs and AI chips in production environments. SDCs originate from a range of low-level causes including manufacturing defects, aging-induced degradation, process variation, particle strikes, and electromagnetic interference. In this work, we revisit the modeling debate between software-level and microarchitecture-level fault injection for estimating SDC vulnerability, in the context of programmable AI accelerators. While software-level (hardware agnostic) techniques are fast and easy to deploy, studies on CPUs and GPUs have shown they produce misleading results due to their lack of the hardware notion which determines faults propagation or filtering. We show that these issues also persist dramatically in AI accelerators. Using detailed microarchitectural modeling, we demonstrate that even so-called hardware-aware software-level approaches can misestimate FIT rates by more than 4 × across realistic accelerator configurations. Our findings support microarchitecture-level simulation as the most effective tradeoff point between accuracy and scalability for early-stage reliability analysis of programmable AI hardware.