Investigating the Robustness of Arabic Offensive Language Transformer-Based Classifiers to Adversarial Attacks
Maged Abdelaty, Shaimaa Lazem · 2024
Text transformer models have proven effective in many downstream tasks, including offensive language detection. However, these models can be vulnerable to adversarial examples, in which the attacker perturbs the original text to cause the model to misclassify. Generating adversarial text examples is challenging due to the text's discrete and semantic nature. In this paper, we use an Explainable AI (XAI)-based approach to generate adversarial examples that evade a transformer model fine-tuned to detect Arabic offensive examples. The results show that by substituting just one word (unigram) while preserving the semantic and syntactic structure of the example, we can achieve up to 30% success rate in fooling the model.