What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
Yijun Yang, Ruiyuan Gao, Yu Li, Qiuxia Lai, Qiang Xu · 2022
Adversarial examples (AEs) pose severe threats to the applications of deep neural networks (DNNs) to safety-critical domains, e.g., autonomous driving. While there has been a vast body of AE defense solutions, to the best of our knowledge, they all suffer from some weaknesses, e.g., defending against only a subset of AEs or causing a relatively high accuracy loss for legitimate inputs. Moreover, most existing solutions cannot defend against adaptive attacks, wherein attackers are knowledgeable about the defense mechanisms and craft AEs accordingly.