Stability Analysis of Deep Neural Networks under Adversarial Attacks and Noise Perturbations
Parisa Eslami, Houbing Herbert Song · 2024
The model’s uncertainty estimation is what adversarial attacks target and exploit. However, the attacker may not always have a clear understanding of how the model estimates uncertainty. Different types of uncertainty can impact these attacks differently because not all sources of uncertainty are equally affected by the attack strategies. We investigate the impact of various noise distributions as well as adversarial attacks on distinct networks. Our objective is to establish a threshold for quantifying network robustness. To achieve this, we analyze three network models, progressively increasing the depth of layers. Our analysis incorporates Shannon entropy and Kullback-Leibler divergence to assess model uncertainty. This directs our attention toward identifying a criterion for measuring uncertainty. Notably, our findings shed light on the manipulation of neural network uncertainty through adversarial attacks, highlighting variations across diverse datasets and models.