Intriguing Properties of Universal Adversarial Triggers for Text Classification
Kexin Zhao, Runqi Sui, Xuedong Wu, Wenchuan Yang · 2024
The universal adversarial trigger (UniTrigger) is an extremely destructive attack method, posing significant threats particularly in the domain of text classification. However, our understanding of UniTrigger, in both attack and defensive contexts, remains considerably limited. In this work, we conduct an extensive experimental study and compared the attack effects of UniTrigger on various text classification models. Our analysis reveals several intriguing properties of UniTrigger: (1) UniTrigger exhibits universality and has a remarkable attack effect in attacking models with biased predictive behaviour, in some cases even reducing their accuracy to 0%. (2) UniTrigger exhibits a vulnerability: its attack efficacy can be significantly reduced by adjusting specific hyperparameters associated with the model’s architecture. e.g., increasing the convolutional kernel size in CNN can reduce the attack success rate by 60% to 90%. (3) UniTrigger can strategically select different concatenation positions based on the targeted models, enhancing the flexibility of the attack. Meanwhile, we find that adversarial attacks can contribute to our understanding of predictive behavior and position information of models, providing a novel perspective in model research.