Adversarial Training and Provable Defenses: Bridging the Gap
Mislav Balunović, Martin Vechev · Repository for Publications and Research Data (ETH Zurich) · 2020
We present COLT, a new method to train neural networks based on a novel combination of adversarial training and provable defenses.The key idea is to model neural network training as a procedure which includes both, the verifier and the adversary.In every iteration, the verifier aims to certify the network using convex relaxation while the adversary tries to find inputs inside that convex relaxation which cause verification to fail.We experimentally show that this training method, named convex layerwise adversarial training (COLT), is promising and achieves the best of both worlds -it produces a state-of-the-art neural network with certified robustness of 60.5% and accuracy of 78.4% on the challenging CIFAR-10 dataset with a 2/255 L ∞ perturbation.This significantly improves over the best concurrent results of 54.0% certified robustness and 71.5% accuracy.Published as a conference paper at ICLR 2020 using adversarial training.Overall, we can see this method as bridging the gap between adversarial training and provable defenses (it can conceptually be instantiated with any convex relaxation).We experimentally show that the method is promising and results in a neural network with state-of-theart 78.4% accuracy and 60.5% certified robustness on the challenging CIFAR-10 dataset with 2/255 L ∞ perturbation (the best known existing results are 71.5% accuracy and 54.0% certified robustness from concurrent work of Zhang et al. ( 2020)). Main Contributions Our key contributions are:• A new method, which we refer to as convex layerwise adversarial training (COLT), that can train provably robust neural networks and conceptually bridges the gap between adversarial training and existing provable defense methods.• Instantiation of convex layerwise adversarial training using linear convex relaxations used in prior work (accomplished by introducing a projection operator).• Experimental results showing convex layerwise adversarial training can train neural network models which achieve both, state-of-the-art accuracy and certified robustness on CIFAR-10 with 2/255 L ∞ perturbation.