Provably Accurate Memory Fault Detection Method for Deep Neural Networks

Omid Aramoon, Gang Qu · 2021

Deep Neural Networks (DNNs) have been widely deployed in real-world systems, many of which have strict safety constraints. Soft errors on memory acceleration platforms for DNNs can degrade their inference accuracy and result in silent data corruption, which can have severe consequences in safety-critical applications. No doubt to say, efficient and effective techniques to detect and mitigate memory faults are needed. In this paper, we propose a novel methodology to diagnose the presence of faults in the memory of DNN accelerators. Our method queries the protected DNN with a set of specially crafted test cases that can accurately reveal if model parameters stored in the hardware are faulty. We provide a theoretical guarantee for the performance of our method and conduct systematic proof-of-concept experiments by simulating memory faults on computer vision models. Our empirical evaluations corroborate the effectiveness and efficiency of our approach. Detecting faults with our method requires simple decision-based access to the inference capability of the DNN, and does not require any additional functionality from the accelerator, which makes our method ideal for legacy systems.

Read the paper · More papers on PaperTik