AIRABIC: Arabic Dataset for Performance Evaluation of AI Detectors
Hamed Alshammari, Ahmed EI-Sayed · 2023
In the rapid expansion of Large Language Models (LLMs), such as ChatGPT, AI-generated text detection models have made substantial advancements, marking meaningful progress in several research and industrial applications. However, the performance of these models in the Semitic languages' context, especially regarding diacritic usage, continues to be inadequately explored. One of these Semitic languages that is still commonly used is Arabic language. To study the performance of the recent AI-generated text detection models on the Arabic language, this paper introduces the AIRABIC dataset, a combination of 1000 examples encompassing 500 human-written passages from 41 unique sources and an equal number of AI-generated texts from the OpenAI's Generative Pre-trained Transformer version 3.5 (GPT-3.5 Turbo ChatGPT). This study focuses on the performance evaluation of two prominent AI-generated text detectors, GPTZero and OpenAI's Text Classifier. Our findings reveal that GPTZero achieves an overall accuracy rate of 62.6%, while the OpenAI Text Classifier exhibits a 50% rate of biased categorizations when analyzing human-written text. Further analysis shows the design gaps of these detectors, especially for identifying human-written Arabic text, mainly when diacritics are involved. The detection accuracy for human-written texts with diacritics is as low as 30% for GPTZero and 0% for the OpenAI Text Classifier. Results also show the potential of diacritics to reduce the detectors' accuracy and the need to handle them in the detectors' design process.