Escaping Adversarial Attacks with Egyptian Mirrors
Olga Saukh · 2023
Adversarial robustness received significant attention over the past years, due to its critical practical role. Complementary to the existing literature on adversarial training, we explore weight-space ensembles of independently trained models. We propose a defense against adversarial examples which takes advantage of the latest empirical findings on linear mode connectivity of overparameterized models modulo permutation invariance. Egyptian Mirrors defense escapes adversarial attacks by moving along linear paths between pair-wise aligned functionally diverse models, while frequently and arbitrary changing ensembling direction. We evaluate the proposed defense using adversarial examples generated by FGSM and PGD attacks and show improvements up to 8% and 33% test accuracy on 2-layer MLP and VGG11 architectures trained on GTSRB and CIFAR10 datasets respectively.