Fine-tuning Language Models for AI vs Human Generated Text detection

Sankalp Bahad, Yash Bhaskar, Parameswari Krishnamurthy · 2024

In this paper, we introduce a machinegenerated text detection system designed to tackle the challenges posed by the proliferation of large language models (LLMs).With the rise of LLMs such as ChatGPT and GPT-4, there is a growing concern regarding the potential misuse of machine-generated content, including misinformation dissemination.Our system addresses this issue by automating the identification of machine-generated text across multiple subtasks: binary human-written vs. machine-generated text classification, multiway machine-generated text classification, and human-machine mixed text detection.We employ the RoBERTa Base model and fine-tune it on a diverse dataset encompassing various domains, languages, and sources.Through rigorous evaluation, we demonstrate the effectiveness of our system in accurately detecting machine-generated text, contributing to efforts aimed at mitigating its potential misuse.

Read the paper · More papers on PaperTik