DeepSHAP Summary for Adversarial Example Detection
Yi-Ching Lin, Fang Yu · 2023
The applications of deep learning are widely used across various fields. Additionally, Explainable AI (XAI) helps interpret model predictions, making them more reliable and trustworthy. Taking advantage of XAI techniques, we propose adopting decision logic from explanations to detect adversarial examples. We show that there exists a different interpretation distribution between normal and adversarial examples, as well as diverse decision logic of networks to distinguish them. Specifically, we first use DeepSHAP values to calculate the neuron contribution of the classification model layer-by-layer and show how to utilize this layer-wise explanation to distinguish between normal and adversarial examples. Second, we select critical neurons according to their SHAP values, generate a bitmap to represent the decision logic that interprets distribution of critical neurons, and propose a new approach that uses the decision logic instead of SHAP signature to detect adversarial examples. The preliminary results against the CIFAR-10 dataset demonstrate that the more SHAP layer information is given, the better accuracy can be achieved. Furthermore, using decision logic that concentrates on critical neurons can 1) outperform single or two layer SHAP signature detection approach, and 2) achieve competitive accuracy to all-layer SHAP signature detection with less resource requirement. We also show that the effectiveness of the proposed approach can be transferred to detect untrained attacks.