ARR: An Attention-Based Robust Residual Network Architecture
Yunlong Li, Qingguo Xu, Shengbo Chen, Fei Zheng, Guoquan Qiang · 2024
As controllers of safety-critical systems, the robustness of feedforward neural networks is crucial. Although current research mainly focuses on developing adversarial training methods, the impact of network structure on robustness cannot be ignored. In this paper, we propose a novel robust deep network, AttentiveRobustResNet(ARR), which integrates the residual network architecture and the multi-head self-attention mechanism to effectively improve the model’s robustness in the face of adversarial samples without increasing additional model parameters. By using the multi-head self-attention mechanism to replace part of the convolution operation for feature extraction, the global information capture ability of the network is improved. By adjusting the sub-components of the network (activation function, data augmentation), the impact of abnormal phenomena such as gradient disappearance and global information loss in adversarial training on the robustness of adversarial training is alleviated. Based on adopting various adversarial training methods, this paper conducts extensive experimental validation of different adversarial attacks on multiple datasets. The results show that AttentiveRobustResNet significantly improves adversarial robustness compared with other neural network models under the same parameter scale. In particular, compared to the baseline network ResNet50, our model achieves an ${8 \%}$ improvement in clean accuracy and a 5% improvement in adversarial robustness (AA) under the TRADES adversarial training framework.