Adversarial Example Detection Using Robustness against Multi-Step Gaussian Noise

Yuto Kanda, Kenji Hatano · 2024

Deep Neural Networks have been wildly successful in various tasks. However, they are known to misclassify against cleverly created input called adversarial examples easily. To realize reliable DNNs, various methods for detecting adversarial examples have been proposed to eliminate adversarial examples in advance. One of them, the input destruction approach, focuses on the difference in robustness between adversarial and clean examples and performs the adversarial example detection based on the robustness of the input image when it is intentionally destroyed by Gaussian noise. Existing detection methods using Gaussian noise use a single, empirically determined Gaussian noise, raising the question of whether or not they can truly maximize anomalies. We propose a novel AE detection method using multiple and multi-step Gaussian noise, which extends the existing single Gaussian noise-based input destruction approach so as to maximize the anomaly of adversarial examples.

Read the paper · More papers on PaperTik