Vanilla Feature Distillation for Improving the Accuracy-Robustness Trade-Off in Adversarial Training

Guodong Cao, Zhibo Wang, Xiaowei Dong, Zhifei Zhang, Hengchang Guo, Zhan Qin, Kui Ren · IEEE Transactions on Dependable and Secure Computing · 2024

Adversarial training has been widely explored for mitigating attacks against deep models. However, a critical limitation of existing works is that robustness enhancement is at the cost of noticeable accuracy degradation. To achieve a better trade-off between robustness and accuracy, we propose the Vanilla Feature Distillation Adversarial Training (VFDAT), which conducts knowledge distillation from a pre-trained model (optimized towards high accuracy) to guide adversarial training model towards generating high-quality and well-separable features by constraining the obtained features of natural and adversarial examples. More specifically, both adversarial examples and their natural counterparts are forced to be aligned in feature space by distilling predictive representations from a pre-trained natural model. In this way, the adversarial training model can be updated towards maximally preserving the accuracy as gaining robustness. A key advantage of our method is that it can be universally adapted to and boost existing works. Exhaustive experiments on various datasets, classification models, and adversarial training algorithms demonstrate the effectiveness of our proposed method.

Read the paper · More papers on PaperTik