Human or GenAI? Characterizing the Linguistic Differences between Human-Written and LLM-Generated Text
Brook Zeleke, Amish Soni, Lydia Manikonda · 2025
The main goal of this paper is to investigate if we can identify and characterize how the text generated by Generative AI (genAI) systems is different from the text written by humans.To achieve this goal, our study uses a publicly available dataset curated from a popular subreddit -\r\eli5 -"Explain Like I'm Five" where, the goal is to respond to the questions posed on the community with layperson-friendly explanations.We collect the top-voted answers from the forum as the human-written responses and prompt Ope-nAI's ChatGPT to generate responses to the same set of questions to investigate their similarities and differences.With the help of expert coders and Natural Language Processing approaches, we evaluate how the texts are similar and different.Our results highlight that human responses are typically shorter in length, informal, uses analogies heavily for explanations, and tend to have conclusive answers.Responses of GenAI are longer in length, formal, cites existing law or policies as examples for explanations, and less likely to reach conclusions.Additionally, through an experiment we found that humans have the innate ability to differentiate and identify which text was written by humans or generated by genAI.These preliminary results indicate a promising direction that it is possible to develop and deploy automated approaches to detect if a particular text was written by humans or generated using LLMs and the importance of prompt in generating appropriate responses.