Text Authorship Attribution: Stylometric Insights into Human and LLM-Generated Text
Shifali Agrahari, Samridhi Bisht, Sanasam Ranbir Singh · 2024
The widespread use of large language models (LLMs) has raised concerns about the potential misuse of AI-generated text for deceptive purposes, such as disinformation and spam. While there is research on detecting human and AI-generated text, specific investigations into the differences between texts generated by individual LLMs remain limited. This study aims to explore the distinctions between human-written text and text produced by various LLMs, including Gemini, GPT-Neo, Falcon, LLaMA, and Bloom. Through a comprehensive analysis of lexical and stylistic features, we have found that LLM-generated text tends to be longer, more structured, and less lexically diverse than human-written content. Cross-model classification experiments revealed that although models like Gemini and ChatGPT closely align with human text, detecting AI-generated content across different models remains a challenge. Our findings underscore the need for better detection techniques to effectively distinguish between human and AI-generated text.