Adversarial defense by restricting in-variance and co-variance of representations

Jeevithan Alagurajah, Chee‐Hung Henry Chu · Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing · 2022

Despite high accuracies achieved by deep neural networks (DNNs) in image classification, DNNs have been shown to be highly vulnerable to structured and unstructured perturbations to the input images. Robustness of many existing defense methods for these models suffers greatly when an attacker has full knowledge of the model and can iterate over the model to craft stronger attacks, which is known as white box attacks. We conduct empirical analysis on the representation of DNN under state-of-the-art attacks to find this causes instances to move closer to a false class in representation space when such perturbation is added to the input. This causes the model to make incorrect decisions even when the adversary and clean images are indistinguishable to human perception. Motivated by this observation, we propose a class-wise disentanglement on intermediate representations of DNN. Specifically, we force DNNs to learn same-class representations to be closer and different-class representations to be maximally farther apart. Moreover, we force the representations of clean and noisy data to be closer if it comes from the same class by restricting its variance in representation. In this approach, a DNN is forced to learn decision boundaries that are distinct for each class with clear separation. We observe that this constraint on representations enhance the robustness of learned models even against strongest white-box attacks. Further we evaluate extensively on both white-box and black-box settings and show significant gains in comparison to state-of-the art defenses. (Implementation:https://github.com/Jeevi10/AICR)

Read the paper · More papers on PaperTik