The Precision of Model to Detect AI-generated Text: Comparison on AI Text Detectors and Impact of Paraphrasing
Andrew Nicholas Jansen, Jacky Suwandy, Yohan Muliono, Nadia Nadia, Franz Adeta · 2025
The swift advancement of generative technologies such as LLMs has resulted in a variety of new challenges, especially in detecting AI-generated text. This study will evaluate 10 different AI models in distinguishing Indonesian AI-generated and human-generated text and the impact of paraphrasing on model detection outcomes. Each model will be trained using a dataset consisting of 224 paragraphs divided evenly between AI and human writing with each paragraph containing approximately 100-200 words. The models are validated using a paraphrased version of all 224 training dataset paragraphs, generated by Quillbot. The evaluation results show that partially paraphrased human-generated text and AI-generated text continue to be classified as their original categories. Three models in particular — indobenchmark/indobert-base-p1, indobenchmark/indobert-base-p2, and indobenchmark/indobert-large-p2 — demonstrate high accuracy in detecting the validation dataset. While all models can generally differentiate between Indonesian AI-generated and human-generated text, they still face challenges when dealing with paraphraser, especially paraphrased human-generated text.