An Empirical Evaluation of Adversarial Examples Defences, Combinations and Robustness Scores

Aleksandar Janković, Rudolf L. Mayer · 2022

Over the past few years, deep learning has been dominating the field of machine learning in applications such as speech, image, and text recognition, which lead to an increased use of deep learning techniques in safety-critical tasks. However, Neural Networks are vulnerable to adversarial examples, i.e. well-crafted small perturbations of the input that aim to disturb the prediction correctness. Therefore, robustness and security of deep learning models has become a major concern, indirectly also affecting safety. In this paper, we therefore evaluate several state-of-the-art white- and black-box adversarial attacks against Convolutional Neural Networks for image recognition, for various attack targets. Further, defences such as adversarial training and pre-processors are evaluated. Moreover, we investigate whether combinations of them can improve these defences. Finally, we examine whether attack-agnostic robustness scores such as CLEVER are able to correctly estimate the robustness against our large range of attack. Our results indicate that pre-processors are very effective against attacks with adversarial examples that are very close to the original images, that combinations can improve the defence strength, and that CLEVER is insufficient as the sole indicator of robustness.

Read the paper · More papers on PaperTik