A widely used Generative-AI detector yields zero false positives

Samuel D. Gosling, Kate Ybarra, Sara K. Angulo · Aloma: revista de psicologia, ciències de l'educació i de l'esport Blanquerna · 2024

The widespread availability of generative-AI using Large Language Models (LLMs) has provided the means for students and others to easily cheat on written assignments – that is, students can use AI to generate text and then submit that work as their own. A variety of technical solutions have been developed to detect such cheating. However, concerns have been raised about the dangers of falsely identifying real students’ responses as having been generated by AI. Here we evaluate a Generative AI detector that comes as an option with Turnitin, a widely used plagiarism-detection platform already in use at many universities. We compare 160 responses written by students in a class assignment with 160 responses generated by ChatGPT instructed to complete the same assignment. The ChatGPT responses were generated by 16 different prompts crafted to mimic those that plausibly might be given by individuals seeking to cheat on an assignment. The AI scores for the AI generated responses were significantly higher than the AI scores for the human-generated responses, which were all zero. Clearly, an arms race is set to develop between technology that facilitates cheating and technology that detects it. However, the present findings demonstrate that it is at least possible to deploy technical solutions in this context. Looking ahead, as various AI-methods become legitimate tools and are more seamlessly integrated into almost every aspect of daily life, it is unlikely that purely technical solutions will suffice. Instead, guidelines surrounding academic integrity will have to adapt to new conceptualizations of academic mastery and creative output.

Read the paper · More papers on PaperTik