Adversarial Attacks on Deep Neural Networks using Generative Perturbation Networks
Sudharman K. Jayaweera · 2023
A new deep learning (DL) network called Generative Perturbation Network (GPN) is proposed. The GPNs are capable of learning to modify inputs to deep neural network (DNN) models trained to classify images in imperceptible ways leading to misclassification. Unlike previous approaches to generate such adversarial samples to DNN classifiers, the proposed GPN does not need to know the architecture of the neural network that it is trying to attack nor have access to the data used to trained it. It is shown that with just being able to observe the final output labels from the trained classifier to any given input image, the GPNs are able to learn to minimally perturb the images to achieve misclassification. Simulation results show that a proposed GPN can easily degrade the 98.65% accuracy of a trained CNN on the MNIST hand-written digit dataset down to about 52%. Simulation results also show that different regularizations can be used in the generator loss function to achieve various visually desirable characteristics in the generated adversarially perturbed images.