Investigating the Limitations of Adversarial Training for Language Models in Realistic Spam Filter Deployment Scenarios
Luay Albtosh · Advances in information security, privacy, and ethics book series · 2024
Adversarial training has emerged as the prevailing approach for fortifying natural language processing (NLP) models against adversarial attacks, bolstering their accuracy and overall performance. However, little research has investigated the long-term viability of adversarial training. Upon deployment, models are often updated with fresh, non-adversarial data samples. This research seeks to systematically assess the longitudinal effects of adversarial training on language models, particularly as they undergo organic evolution over periods of time, emphasizing the task of spam detection. Extensive experiments are conducted using multiple spam classification models on various benchmark dataset. The findings reveal that the effectiveness of adversarial training is contingent upon the task and dataset. When trained on a consistent dataset, models often exhibit commendable predictive accuracy. However, their efficacy tends to wane when subjected to novel datasets, a trend observed.