Improved Understanding, Detection, and Diagnosis of Faults in Deep Neural Networks

Sigma Jahan · DalSpace (Dalhousie University) · 2026

Deep Neural Networks (DNNs) are central to many critical AI applications, yet ensuring their reliability remains difficult. Faults in DNNs may originate from training data, model architecture, learning dynamics, or hardware-software interactions. Such faults often produce no explicit errors, yet degrade model performance or cause unexpected behaviours. A faulty attention mechanism, for example, may not trigger a crash while still causing poor generalization or incorrect outputs. This thesis addresses the understanding, detection, and diagnosis of faults across neural architectures of increasing complexity (e.g., FFNNs to transformers) through five studies. First, we study 5,278 faults and show that established bug localization techniques (e.g., IR/Spectrum-based) perform 27%-34% worse in mean average precision on deep-learning faults than on traditional software faults. Second, we introduce DEFault, a hierarchical technique for detecting, categorizing, and diagnosing faults in earlier DNN programs built with FFNNs, CNNs, RNNs, and LSTMs. DEFault combines static and dynamic analyses and improves fault diagnosis by 11.54% over prior baselines on real-world faulty DNN programs. Third, we characterize attention-specific faults through a formal taxonomy derived from 555 real-world cases, identifying seven fault categories and 25 root causes. Fourth, we provide Hessian-based evidence that attention faults can propagate through cross-parameter interactions that first-order methods (i.e., gradients) may not reveal. Fifth, we develop DEFault++, a three-level hierarchical technique for detecting, categorizing, and diagnosing faults in transformer models. We construct a benchmark of 5,556 labelled instances covering 12 fault categories and up to 45 root causes. DEFault++ uses this benchmark to learn fault-propagation patterns and predict both fault category and root cause. A developer study with 21 practitioners shows that DEFault++ improves repair-action accuracy from 57.1% to 83.3%. Together, these studies connect the behaviour of DNN faults with new techniques for detecting and diagnosing them.

Read the paper · More papers on PaperTik