Research on Detecting ChatGPT Rewritten Text

Zhihui Wang, De Li, Xun Jin · 2024

ChatGPT enables users to efficiently access information services, driving AI-generated content into the public eye. Generated content is widely used in various fields, including education, entertainment, and medicine. The advancement of models like GPT poses significant challenges in identifying the authors of texts. ChatGPT can mimic human dialogue and provide users with information and advice, but the reliability and credibility of the information it provides have not been verified. For individuals who are not computer professionals, it is difficult to distinguish these high-quality texts from those written by humans. In addition to academic research, commercial organizations have also attempted to classify text sources and have launched online detection tools. Although preliminary research has been carried out on the topic of identifying text generated by large language models, the complexity of the vocabulary and the expressiveness of the generated text further complicate the detection process. The current methods do not ensure robustness and transferability, and the difficulty of detecting text generated by different prompts varies, particularly when the data is embellished with GPT. Although there is no standardized measurement tool, the disparity between the two texts is decreasing. In this paper, we propose a method that utilizes text features as auxiliary signals and expands upon them using the pre-trained language model Roberta. This method has demonstrated exceptional performance on datasets that have been modified by ChatGPT.

Read the paper · More papers on PaperTik