Detection of AI-Generated Text
Rakhi Bharadwaj, Raj Dugad, Akshata Gile, Gyaneshwari Patil, Shrinivas Hatyalikar · 2024
This research paper investigates the challenges posed by the proliferation of AI-generated content in the context of plagiarism detection and classification. Leveraging advanced language models, particularly the GPT-2 model, in tandem with tailored metrics and confidence thresholds, our approach provides a comprehensive solution to tackle this pressing issue. By employing tokenization and perplexity calculation techniques, we evaluate the coherence and naturalness of text inputs, thereby identifying potential AI-generated content. Additionally, burstiness score computation aids in discerning repetitive language patterns, further contributing to the detection process. Furthermore, integrating a BERT-based text classification model enables the classification of text based on its conformity to AI-generated characteristics, with a set confidence threshold determining the likelihood of AI generation. Through experimental evaluation, our approach demonstrates high accuracy in detecting AI-generated content while minimizing false positives. Moreover, we explore the impact of adjusting the confidence threshold on the sensitivity and specificity of the classification mechanism. Overall, this research presents a robust framework for AI plagiarism detection and classification, offering a scalable and adaptable solution for effectively identifying AI-generated text.