On Detecting and Defending AdvDrop Adversarial Attacks by Image Blurring and Adaptive Noising
Hequan Zhang, Song Chang Jin · 2023
AdvDrop, a recently emerged adversarial attack method, generates adversarial samples by dropping information from the clean images. Compared with the traditional adversarial attack based on adding perturbation, AdvDrop is more robust against current de-noising based defense techniques, which poses a great threat to the security of DNNs. To defend AdvDrop, this paper proposes an image blurring detection method to detect AdvDrop adversarial samples and also proposes a defense method based on adaptive noising combined with fast and flexible denoising convolution neural network (FFDNet) filtering to restore the correct class of adversarial samples. Image blurring is first conducted to add a mask to the spectrum of the input image in the frequency domain. Thanks to the robustness of the DNN, the blurring operation has no significant impact on the accuracy of the clean samples but prevents the attackers from extracting useful image features. As a result, the adversarial samples can be detected by analyzing the consistency of the predicted classification output by DNN. In the defense method, we introduce adaptive noising to enhance the high-frequency information in each image. Combined with FFDNet filtering, this strategy increases the diversity of local changes in the image while preserving the overall image structure. As a result, DNN can focus more on essential image features rather than those created by the attacker. We evaluated our detection and defense methods on Cifar-10 and ImageNet datasets for some DNN models, and the results surpassed the current state-of-the-art methods..