Interpretation Of White Box Adversarial Attacks On Machine Learning Model Using Grad-CAM

Uppalapati Hari Krishna Sai, Vinay Sai Yogeesh, N. D. Vindya, Akanksha P Mulgund, Bhaskarjyoti Das · 2024

ML models are now finding extensive applications in sensitive fields such as healthcare. These models contribute significantly to diagnosis and decision-making. However, they suffer from adversarial attacks wherein small, often imperceptible, changes in input data produce undesirable outputs. Such vulnerabilities can further threaten patient safety and privacy. We conduct this study to interpret adversarial attack impact in CNN model using Grad-CAM. Grad-CAM is a tool that visualizes the model’s decision making through highlighting parts of an input image that influence its predictions. We have used SqueezeNet model, a lightweight convolutional network for the classification of ophthalmic images which includes 4 classes - normal, cataract, diabetic retinopathy, and glaucoma. We then subjected our model to four standard attacks: the Fast Gradient Sign Method, Projected Gradient Descent, Basic Iterative method and Carlini Wagner attack. Our experiments show that each attack heavily distorts the attentions induced in the model by Grad-CAM, and it, in turn, produces incorrect classifications. Therefore, such findings emphasize the need for developing stronger ML models as well as the utility of interpretability tools in explaining and countering adversarial threats. Future work will comprise integration of insights provided by Grad-CAM together with better detection systems to enhance the resilience of the model.

Read the paper · More papers on PaperTik