Generating Transferable Perturbations via Intermediate Feature Attention Weighting
Qi Wang, Jie Jiang, Yishan Li · 2023
Deep neural networks have achieved significant success across multiple domains. Nonetheless, owing to their "black box" nature, the internal mechanisms and inference outcomes of deep neural networks frequently face scrutiny. Consequently, prior to deploying deep learning models, assessing their security is essential, particularly in specific domains with stringent security requirements. In recent years, black-box attacks based on adversarial sample transferability have garnered increasing attention due to their relevance to real-world situations. Attackers must generate adversarial samples on local substitute white-box models and input them into other unknown black-box models for attack. However, adversarial samples tend to fall into local optima and overfit the substitute models, resulting in limited transfer attack capabilities. We compute attention weights to reflect the significance of feature maps and guide the generated perturbations to disrupt the shared salient features utilized for decision-making in different models, thereby enhancing transferability. By incorporating these eigenvectors with the original features in the loss function, we optimize perturbations to improve the generalization capability of adversarial samples. Lastly, in contrast to the conventional multi-step iterative approaches, we utilize a generator framework for perturbation generation. After training, numerous adversarial samples can be swiftly produced within a minimal timeframe. Our method’s effectiveness and efficiency are validated through extensive experiments.