Adversarial Training for Probabilistic Robustness
Y N Zhang, Yu Chen, Zhen Chen, Wenjie Ruan, Xiaowei Huang, Siddartha Khastgir, Xingyu Zhao · 2025
Deep learning (DL) has shown transformative potential across industries, yet its sensitivity to adversarial examples (AEs) limits its reliability and broader deployment. Research on DL robustness has developed various techniques, with adversarial training (AT) established as a leading approach to counter AEs. Traditional AT focuses on worst-case robustness (WCR), but recent work has introduced probabilistic robustness (PR), which evaluates the likelihood of AEs within a local perturbation range, providing an overall assessment of the model's robustness and acknowledging residual risks that are more practical to manage. However, existing AT methods are fundamentally designed to improve WCR, and no dedicated methods currently target PR. To bridge this gap, we formulate a new min-max optimization as the theoretical foundation for PR-focused AT, and introduce an AT-PR training scheme with numerical algorithms to solve the new optimization problem. Our experiments, based on 70 DL models trained on common datasets and diverse architectures, demonstrate that: i) AT-PR achieves higher improvements in PR than AT-WCR methods; ii) it shows more consistent effectiveness across varying local inputs; iii) it exhibits a reduced trade-off in model's generalization.