Character-level Adversarial Attacks Evaluation for AraBERT’s

Saja Nakhleh, Malik Qasaimeh, Ammar Qasaimeh · 2024

Research has demonstrated the continuous growth of Large Language Models (LLMs) to adversarial attacks, which involve the crafted input samples that can deceive even well-performing models. Minor perturbations to the inputs can lead these models to make incorrect predictions. This vulnerability has been identified in various domains, including computer vision, speech recognition, and natural language processing (NLP). This study investigates the robustness of AraBERT LLMs against char-level Black-Box attacks induced by spelling errors. Employing the Chain-of-Thought prompting technique, we generate adversarial samples to assess the resilience of a fine-tuned AraBERT model and AraBERT v2.0 LLM. Results indicate greater robustness for the fine-tuned model, with a 12% accuracy decrease under the replacement attack. Moreover, longer adversarial samples exhibit increased resilience, with only a 2% accuracy decrease observed in such instances. Conversely, AraBERT v2.0 LLM displays the lowest robustness, experiencing a significant 46% accuracy decrease when subjected to the replacement attack.

Read the paper · More papers on PaperTik