Evaluation of interpretability methods for adversarial robustness on real-world datasets
Anna Chistyakova, Maria Cherepnina, Konstantin Vladimirovich Arkhipenko, Sergey Dmitrievich Kuznetsov, Chang-Seok Oh, Sebeom Park · 2021
Adversarial training is considered the most powerful approach for robustness against attacks on deep neural networks involving adversarial examples. However, recent works have shown that the similar robustness level can be achieved by other means, namely interpretability-based regularization. We evaluate these interpretability-based approaches on real-world ResNet models trained on CIFAR-10 and ImageNet datasets. Our results show that interpretability can marginally improve robustness when combined with adversarial training, however, they bring additional computational complexity making these approaches questionable for such models and datasets.