ArabicTextFool: A Framework for Arabic Adversarial Attacks
Yasmeen Alslman, Arafat Awajan · 2025
In recent years, researchers have increasingly utilized deep learning models due to their enhanced performance in natural language processing applications. Text classification tasks, such as sentiment analysis, have particularly benefited from deep learning models. However, the inherent black-box nature of deep learning raises concerns, as adversaries may exploit vulnerabilities in these models to undermine their effectiveness. This paper introduces a new adversarial attack model called ArabicTextFool, aimed to exploit the vulnerabilities in a version BERT transformer, specifically DistilBERT, which is known for its black-box characteristics. Adversarial samples were generated using a genetic algorithm to identify optimal positions for adding dots/diacritics, enabling the manipulation of the sentence's label while maintaining readability for human audiences. The results demonstrated that the attack had a success rate of 90% in the worst-case scenario. In conclusion, the high success rate of ArabicTextFool underscores the need for defensive mechanisms when employing deep learning models in sentiment analysis.