PEGAT: Prediction Error-Guided Adversarial Training to Enhance Robustness of Deep Learning Models in Autonomous Vehicles

Manzoor Hussain, Zhengyu Shang, Ahmed Dawod Mohammed Ibrahum, Jang‐Eui Hong · IEEE Access · 2025

Adversarial training is a widely used method to improve the robustness of deep learning models in various applications. Although adversarial training enhances the robustness of the target model, it also suffers from an accuracy versus robustness trade-off, meaning that while improving robustness, the accuracy of the target model decreases. Moreover, they also suffer from catastrophic overfitting, where these methods offer reasonable robustness against specific adversarial attacks, but they become ineffective against others. Thus, to solve these issues, we proposed Prediction Error-Guided Adversarial Training (PEGAT). In this method, we first evaluate the normally trained target models against adversarial samples and then calculate class-wise prediction error in terms of loss in the F1-score, which we refer to as PEGAT scores. In the second phase, use the PEGAT scores as guidance to assign weights during the adversarial training process. In this phase, PEGAT adaptively assigns higher weights to classes with high PEGAT scores and lower weights to classes with low PEGAT scores. The PEGAT guides the models to focus on those classes where they struggle to maintain a higher F1 score in the first phase. By doing so, not only is the adversarial robustness of the target model increased, but also the issues such as the accuracy versus robustness tradeoff, catastrophic overfitting, and generalizability are addressed. Evaluating the PEGAT on classification models used in intelligent autonomous vehicle perception systems, we found that the technique could improve the robustness of WideResNet18, WideResNet34, and WideResNet50 by 16%, 6%, and 5%, respectively, without suffering from the aforementioned issues.

Read the paper · More papers on PaperTik