Filtering Adversarial Noise with Double Quantization
AprilPyone MaungMaung, Yuma Kinoshita, Hitoshi Kiya · 2019
Despite deep learning being powerful to solve challenging problems, it is vulnerable towards adversarial examples. To defend these adversarial blind spots in the deep learning, researchers have proposed various approaches. However, conventional adversarial training can reduce the accuracy significantly. In this paper, we propose a method to incorporate quantized images in both training and testing to maintain identical accuracy for both normal and adversarial examples. Specifically, the proposed method utilizes dithering during training and dithering and linear quantization as a mean of adversarial filtering during testing. We evaluated the proposed method with a well-known strong first-order adversary and also conducted experiments in different bit depths. The results suggest that the proposed method achieves 87.14% and 85.28% accuracy for 2-bit and 1-bit dithered models for both normal and adversarial tests on the noise level of 8. In addition, due to having identical accuracy for both adversarial and normal tests, the proposed method can detect adversarial examples if the original test dataset is known. The code for the experiments is released on https://github.com/fugokidi/one-bit-quantization.