GAN-Guarded Fine-Tuning: Enhancing Adversarial Robustness in NLP Models
Tarek Ali, Amna Eleyan, Mohammed Al-Khalidi, Tarek Bejaoui · 2025
The increasing occurrence of assaults on Natural Language Processing (NLP) systems has sparked worry about their resilience and trustworthiness, particularly in crucial applications like sentiment analysis, machine translation and natural language deduction. To tackle these weaknesses this research presents GAN-Guarded Fine-Tuning (GAN-FT) an training structure that utilizes Generative Adversarial Networks (GANs) to bolster the resilience of extensive language models such, as BERT and RoBERTa. In contrast, training GAN with feature matching (GAN FW) uses a generative opponent to generate various and logically connected adversarial instances with less computational burden. Regular assessments on datasets such as IMDB, Yelp and SNLI indicate that GAN-FT consistently decreases the effectiveness of attacks while enhancing accuracy and adaptability across domains. This suggests its effectiveness in countering familiar and new word substitution forms of attack. This study highlights the impact of incorporating GANs into training methods and sheds light on understanding models better and making decisions effectively in this field of research. Though there are scalability and quick implementation challenges, in real-time scenarios, GAN-FT is a starting point for looking into adversarial defence mechanisms and broadening its use in more intricate NLP assignments. The research advances the reliability and security aspects of NLP systems, which is a move towards combatting adversarial risks in the ever-changing AI domain.