Deep Learning Model Defense Against Manipulative Adversarial Attacks: An Investigation on Adversarial Robustness Techniques
K. Nirmala Devi, S. Karkuzhali, P. Yamini Devi, A. Hency Juliet, S. Palanivel, Vadlamudi Syambabu · 2025
The security of deep neural networks (DNNs) is a major worry, despite the fact that deep learning has revolutionised several industries due to its fast development. Misclassifications caused by adversarial attacks, which are undetectable modifications of input data, can have catastrophic effects in autonomous driving, healthcare, and security surveillance, among other sensitive applications. In order to protect DNNs from these manipulative adversarial attacks, this work examines adversarial robustness strategies in detail. Three main approaches—adversarial training, input transformation, and gradient masking—are the focus of our classification and evaluation of existing defence measures. The effectiveness of each approach in strengthening the model's resilience is assessed by considering how well it maintains accuracy and efficiency. We also present a benchmark methodology for testing defence efficacy uniformly against different types and intensities of attacks, including FGSM, PGD, and C&W attacks. We also look into new developments in certified robustness, a field that aims to provide formal guarantees of model resilience against adversarial perturbations within a certain bound. While adversarial training does provide the best robustness in both white-box and black-box situations, hybrid approaches that combine input transformation with adversarial training have the potential to reduce computing load while boosting robustness, according to the experimental data. The importance of developing more accurate and scalable defence mechanisms is emphasised by our findings, which draw attention to important trade-offs between model interpretability, computational complexity, and robustness. In its last section, this study suggests avenues for further investigation into adversarial defence, with a focus on finding solutions that are independent of specific models and utilising adaptive learning techniques. In order to protect DNNs from ever-evolving hostile threats, these developments are crucial for their implementation in practical applications.