Lightweight Neural Networks for Adversarial Defense: A Novel NTK-Guided Pruning Approach

Akhila Reddy Yadulla, Bhargavi Konda, Mounica Yenugula, Vinay Kumar Kasula, Sarath Babu Rakki, Rajkumar Banoth · 2025

Self-Supervised Adversarial Training (SSAT) is a widely used adversarial attack defense method that integrates adversarial examples into the training process, effectively enhancing robustness against attacks. However, the robustness of SSAT models often relies on increasing network capacity, leading to a significant enlargement of model size and restricting its usability. A major challenge is to develop a lightweight adversarial defense method that maintains robustness while reducing model capacity. To address this issue, we propose a novel lightweight adversarial attack defense approach based on Neural Tangent Kernel (NTK)-Guided Pruning and Attention-Based Robust Distillation, integrated with Friendly Adversarial Training (FAT). Our method optimizes adversarial robustness by performing layer-wise adaptive NTK-guided pruning on a pre-trained adversarially robust model, followed by data-filtering-based Attention-Based Robust Distillation on the pruned network to retain essential robustness properties. Experimental evaluations on CIFAR-10 and CIFAR-100 datasets demonstrate that under the same FAT adversarial training setting, our proposed NTK-guided pruning method outperforms existing pruning techniques, yielding a more robust network structure across different FLOPs settings. Furthermore, the combination of NTK-guided pruning and Attention-Based Robust Distillation achieves higher adversarial robustness accuracy compared to other robust distillation techniques. These results validate that our approach successfully reduces adversarial training model capacity while improving robustness, making it highly suitable for edge computing environments in the Internet of Things (IoT).

Read the paper · More papers on PaperTik