Toward Robust Learning via Core Feature-Aware Adversarial Training
Fengpeng Li, Kemou Li, Haiwei Wu, Jinyu Tian, Jiantao Zhou · IEEE Transactions on Information Forensics and Security · 2025
Deep neural networks (DNNs) are inherently vulnerable to adversarial examples (AEs), severely deteriorating model performance on various tasks. Adversarial training (AT) is one of the most effective approaches to enhance model robustness by incorporating AEs into the training process. Notwithstanding the efficacy of AT, recent studies have unveiled that adversarial perturbations on AEs predominantly impact core features—essential for accurate predictions—more than spurious features, which are incidentally aligned with training labels but irrelevant to the model’s classification. This unequal impact induces the models trained with AT to excessively rely on spurious features, resulting in a pronouncedfeature shiftthat compromises robustness and generalization against AEs at inference. In this work, we introduce a novelCore Feature-aware Adversarial Training(COFAT) framework to cope with these challenges. COFAT employscore feature extractionto dynamically generatecore partnersby selectively retaining benign sample regions on feature maps with high-weight while masking low-weight ones, thereby ensuring the model focuses on core features. Furthermore,contrastive feature alignmentis proposed to reduce intra-class feature distances and increase inter-class separability by maintaining a center bank of class feature representations, thus mitigating reliance on spurious features. Compared to state-of-the-art AT methods, COFAT demonstrates superior performance against diverse adversarial attacks. Remarkably, COFAT improves the robustness of ResNet-18 against AutoAttack on CIFAR-10, SVHN, CIFAR-100, and Tiny ImageNet by approximately 2.14%, 3.20%, 1.69%, and 1.86%, respectively, embodying significant advancements in AT. Our code is publicized at https://github.com/Feng-peng-Li/CoFAT.