Enhancing Model Robustness and Accuracy via Learnable Adversarial Training
Chunlong Fan, Wanyan Guo, Li Yuan Xu · 2025
In recent years, deep learning models have made significant advancements in enhancing robustness against single-perturbation adversarial attacks, such as$\ell_{p}$-norm attacks. However, the development of defense mechanisms for composite attacks involving multiple semantic perturbations remains a challenge. In this paper, we propose a method that combines projected gradient descent (PGD) with sequential semantic perturbations to generate composite adversarial examples (CAEs), providing a more comprehensive evaluation of model robustness in various scenarios. Existing adversarial training methods primarily focus on improving robustness against single-type attacks. We introduce learnable adversarial training (LAT), which leverages the classification boundaries of clean models to guide the training of robust models. In contrast, our approach not only enhances the model's defense against composite perturbations but also significantly reduces the loss of natural accuracy. Experiments on the CIFAR-10, CIFAR-100 and Tiny ImageNet datasets show that our proposed training method outperforms traditional$\ell_{\infty}$-norm boundary-based adversarial training, demonstrating superior robustness against various attack types. In summary, our method strikes an optimal balance between defense against composite perturbations and natural accuracy, showing strong potential for practical applications.