Harnessing the Power of the GPT Model to Generate Adversarial Examples

Rebet Keith Jones, Marwan Omar, Derek Mohammed · 2023

In this paper, we propose a method for generating adversarial examples in the text domain using GPT-2, a state-of-the-art language model. Our method employs an iterative algorithm to produce perturbations to input text samples, creating adversarial examples capable of fooling sentiment analysis models. We evaluate our approach on three widely-used benchmark datasets for sentiment analysis: Yelp, MR, and IMDB. Our results show that our approach can generate highly effective adversarial examples that significantly degrade the performance of sentiment analysis models. Specifically, we achieved a decrease in accuracy of up to 67.3% on the Yelp dataset, 68.1% on the MR dataset, and 52.5% on the IMDB dataset. We also discuss the limitations of our approach and the open challenges in this field. Overall, our study demonstrates the potential of GPT-2 for generating effective adversarial examples in natural language processing tasks.

Read the paper · More papers on PaperTik