A Comprehensive Review of Adversarial Learning and Impact of Unsharing Weights Across Classes
S. Vignesh, Bhawna Bhawna, Shubham Anand, Shailender Kumar · 2021
It is evident from recent developments that deep learning has the potential to be a fundamental part in almost every new technology that comes in the future. Deep learning models have performed exceedingly well on standard image classification problems, but their performance drops drastically when presented with adversarial inputs that are created by adding specific small perturbations to the original image. This paper will be divided into two sections. The first section will be a complete review of the existing research in this field. We will provide the reader with the basic concepts of adversarial learning and a broad classification of various adversarial attacks and defenses. In the second section, we propose an ensemble model with max voting and test the impact of adversarial attacks by converting the 10-class problem over MNIST images into 10 binary classification problems. Each weight in the middle layers of a multi-class neural network is shared across all the output classes of the model. We call them the shared weights of the network. Our proposed model consists of no shared weights, shows a slight improve in accuracy against adversarial samples and can detect out-of-domain inputs.