Detecting Auto-generated Texts with Language Model and Attacking the Detector

Saint Petersburg, Russia, Mikhail Orzhenovskii · Computational Linguistics and Intellectual Technologies · 2022

We propose a simple approach to the detection of automatically generated texts. A pre-trained language model, fine-tuned on the shared task’s dataset, achieved 3rd place on the binary task leaderboard with 82.6% accuracy. In the multi-task leaderboard, the language model achieved an F1 score of 64.5% after being fine-tuned with the same procedure. In order to investigate the weaknesses of this approach, we explore two possible attacks on the detector: selecting from language model outputs and directed beam search. These attacks reduce the likelihood of detecting the generated texts without significant loss in quality. Both attacks do not require retraining the generative model and are applied at inference time.

Read the paper · More papers on PaperTik