Fight Perturbations With Perturbations: Defending Adversarial Attacks via Neuron Influence
Ruoxi Chen, Haibo Jin, Haibin Zheng, Jinyin Chen, Zhenguang Liu · IEEE Transactions on Dependable and Secure Computing · 2024
The vulnerabilities of deep learning models towards adversarial attacks have attracted increasing attention, especially when models are deployed in security-critical domains. Numerous defense methods, including reactive and proactive ones, have been proposed for model robustness improvement. Reactive defenses, such as conducting transformations to remove perturbations, usually fail to handle large perturbations. The proactive defenses that involve retraining, suffer from the attack dependency and high computation cost. In this article, we consider defense methods from the general effect of adversarial attacks that take on neurons inside the model. We introduce the concept of neuron influence, which can quantitatively measure neurons’ contribution to correct classification. Then, we observe that almost all attacks fool the model by suppressing neurons with larger influence and enhancing those with smaller influence. Based on this, we proposeNeuron-level Inverse Perturbation(NIP), a novel defense against general adversarial attacks. It calculates neuron influence from benign examples and then modifies input examples by generating inverse perturbations that can in turn strengthen neurons with larger influence and weaken those with smaller influence. Extensive experiments on benchmark datasets and models show that NIP outperforms the state-of-the-art methods in terms of i)effective- it shows better defense success rate ($\sim \!\!\times 1.45$) against 13 adversarial attacks; ii)elastic- it maintains better defense ($\sim \!\!\times 3.4$in the worst case) on large perturbations; iii)efficient- it runs with only$\sim \!\!1/6$time cost; iv)extensible- it can be applied to speaker recognition models and Baidu online image platforms. We further evaluate NIP against potential adaptive attacks and provide interpretable analysis for its effectiveness.