Deep Learning Strategies for Identifying Machine-Generated Text

Annepaka Yadagiri, Partha Pakray · 2025

Generative AIs like LLMs are now accessible to the general public. For example, students can utilize these tools to create essays or complete theses. However, how is a teacher supposed to determine if a text was composed by the student or an AI? Using deep learning techniques, we investigate novel and classic approaches for detecting text created by artificial intelligence. We also study the more complex instance when the AI is asked to write the text in a way that a human would not recognize as AI-generated, as we discovered that categorization is more challenging in this scenario. For our studies, we used llm-detect-ai-generated text from the Kaggle competition dataset, which included texts written by students and texts produced using different LLMs. Our top systems achieve an accuracy of 0.98% and F1 scores of more than 0.98% in classifying simple and complex texts produced by humans and AI-Genereted. The systems combine features such as TF-IDF vectorization, word2Vec, and word embedding properties. Our findings demonstrate that these additional characteristics significantly enhance the performance of several classifiers. Compared to deep learning models, our best-performing model for detecting AI-generated text outperforms even the fine-tuned ROBERTA-Open-AI classifier, achieving an accuracy of 0.98%. This underscores the efficacy of our proposed approach in distinguishing between human and AI-generated content.

Read the paper · More papers on PaperTik