RANDOM ENSEMBLE OF LOCALLY OPTIMUM DETECTORS FOR DETECTION OF ADVERSARIAL EXAMPLES
Amish Goel, Pierre Moulin · 2018
Deep neural networks achieve state-of-the-art performance for several image classification problems but have been shown to be easily fooled by adversarial perturbations which slightly modify a legitimate image in a specific direction and are visually indistinguishable from the original. This presents a security risk for applications such as autonomous systems. We tackle the problem of detecting such "forgeries" using a locally optimal detector which is well suited to detecting weak signal perturbations. We present a procedure for learning the forgery detector from a training set, using Gaussian Mixture Models (GMM) for modeling image patches. A random ensemble of patches is used for detection of the forgery. The reliability of our forgery detector is assessed for several image classification tasks.