Comparing Fine Tuned-LMs for Detecting LLM-Generated Text

Sanidhya Madhav Shukla, Chandni Magoo, Puneet Garg · 2024

With the introduction of LLMs (Large Language Models), NLP has demonstrated strong skills in a variety of applications, including question answering, language translation, and text summarization. As a result, their use in our day-to-day activities and at work has gained general acceptability. It is nevertheless imperative to develop cutting-edge procedures or approaches since, despite their enormous capabilities, they are still unable to live up to human expectations. It is crucial to apply the cutting-edge methods that enable LLMs to specialize in specific fields and reduce any potential for abuse. In contrast, our study has concentrated on employing fine-tuned LLM as detectors to determine whether a text was produced by a machine or by a person. Comparing how well various fine-tuned LMs function as content detectors-that is, as indicators of whether or not material is generated by LLMs-is the goal of this research. Three distinct fine-tuned LLMs are used to solve this binary classification problem: Fas'I'Text, also known as Generative Pre- Trained Transformers; DistilBERT; and BERT, or Bidirectional Encoder Representations from Transformers. The models were tested using various performance metrics, including precision, recall, Fl score, MCC, NPV, and FDR. The results showed that the BERT model outperformed the FasTText and DistilBERT models, which displayed overfitting tendencies when fine-tuned under comparable settings.

Read the paper · More papers on PaperTik