Adversarial Attack Defense Techniques: A Study of Defensive Distillation and Adversarial Re-Training on CIFAR-10 and MNIST

Tahir Elgamrani, Reda Elgaf, Yousra Chtouki · 2024

Adversarial attacks pose a significant challenge to the reliability of machine learning models by introducing imperceptible perturbations that lead to misclassification. This study evaluates the effectiveness of adversarial Re-training on an MNIST model, which achieved a clean accuracy of 99.5% and an adversarial accuracy of 90%, as well as the effectiveness of Defensive Distillation on a CIFAR-10 model, achieving a clean accuracy of 98.83% and an adversarial accuracy of 81.3% against PGD attacks. These results demonstrate that adversarial retraining with FGSM effectively improves robustness for simpler datasets like MNIST, achieving 90% adversarial accuracy post-training. In contrast, Defensive Distillation shows promise for complex datasets like CIFAR-10 due to its ability to generate robust decision boundaries. Compared to prior works that reported varying levels of success for Defensive Distillation across different domains, our results align with findings that dataset complexity significantly influences the efficacy of defense strategies. This analysis underscores the necessity of tailoring defense techniques to dataset characteristics and attack types. The potential impacts of these techniques on end users include enhanced security and reliability in applications that rely on machine learning models, such as biometric authentication, autonomous systems, and financial fraud detection. Defensive strategies like FGSM retraining and Defensive Distillation can ensure more consistent performance under adversarial conditions, reducing vulnerabil-ities and instilling user confidence in AI-driven systems across both simple and complex domains.

Read the paper · More papers on PaperTik