Rethinking Transferable Adversarial Attacks With Double Adversarial Neuron Attribution

Zhiyu Zhu, Zhibo Jin, Xinyi Wang, Jiayu Zhang, Huaming Chen, Kim‐Kwang Raymond Choo · IEEE Transactions on Artificial Intelligence · 2024

Transferable adversarial attacks are a threat to deep neural networks, in particular for black-box scenarios where access to model information is limited. One can, for example, exploit the intermediate layer neurons to generate transferable adversarial samples. However, current works show limitations in providing feature-level attack mechanisms across multiple victim models. In light of the attribution methods, in this paper, we investigate the attribution similarity across different models. We leverage the similarity to incorporate different attribution properties to enhance the sample transferability, for the first time, formulating a novel neuron attribution-based transferable attack termed DANAA++. Specifically, we utilise a range of adversarial attack methods to generate different baseline points through adversarial training. The attribution results are thus obtained along both linear and nonlinear integration paths. In our experiments, the baseline points and integration paths significantly help improve the transferability of adversarial samples. Our approach provides novel insights for building effective attribution-based feature-level adversarial attacks. We release our code at:https://github.com/LMBTough/DANAAPP

Read the paper · More papers on PaperTik