GAN-Based Defense Mechanism Against Word-Level Attacks for Arabic Transformer-Based Model

Hanin Alshalan, Banafsheh Rekabdar · 2024

Recent studies have revealed that deep learning models are vulnerable to adversarial examples, which are small perturbations that are added to the input to fool the model. However, many researchers are proposing new strategies to defend the models that were trained on English corpora. In this paper, we focus on Transformer-based models for low-resource languages like Arabic, which encounter numerous challenges, including word-substitution attacks. These attacks can produce inaccurate and unreliable results and represent a type of adversarial example. To address this, we propose a defense mechanism based on Generative Adversarial Networks (i.e., InfoGANs) that can effectively protect Transformer-based models for Arabic language from word-substitution attacks. Our approach involves applying GAN strategy as defense mechanism by extracting the text embedding vectors from our Transformer-based model (Arabert) and injected them to the generator in the training phase to produce perturbed examples that are difficult to distinguish from the original data, while the discriminator in GAN is tasked with correctly identifying the genuine examples. The proposed method has shown promising results in defending Transformer-based models against word-substitution attacks on the Arabic language. As a result, we find that the accuracy recovered from various types of attacks that we employed on our two datasets ranged from 2% to 20%.

Read the paper · More papers on PaperTik