Human vs AI-Generated Text Detection Using IndoBERT Algorithm for Different Types of Text
Gracia Abadi, Ariana Noya Zaida Amelia, Yohan Muliono, Ika Dyah Agustia Rachmawati, Chrisando Ryan Pardomuan Siahaan, Aditya Kurniawan · 2025
Along with the advancement of Artificial Intelligence (AI) and Large Language Models (LLMs) that has enabled the creation of hyper-realistic, human-like content, it became difficult to tell if a content is created by humans or AI. This difficulty impacts not only how information is absorbed, but also how it is managed and regulated. This research aims to evaluate the effectiveness of IndoBERT-base-p2, a Bidirectional Encoder Representations from Transformers (BERT) model for the Indonesian language, in distinguishing between humanwritten and AI-generated texts and see the accuracy of the model for various informal writing styles and contexts. The type of dataset used in this research is short text i.e. Movie Reviews and Student Experience Surveys. The results show that IndoBERT-base-p2 is highly accurate and precise in distinguishing Human written and AI-generated text with an average accuracy of $\mathbf{9 5, 8 4 \%}$, as well as adapting with different type of datasets used in the experiments.