A Method for Distinguishing Model Generated Text and Human Written Text

Hinari Shimada, Masaomi Kimura · Journal of Advances in Information Technology · 2024

With the rapid development of Large Language Models (LLMs), such as ChatGPT, it is extremely difficult for humans to accurately detect whether sentences are written by LLMs.Especially in academic fields, there is a need to assist human evaluators by discriminating sentences to recognize differences.Assignments such as essays and theses typically require human authors to write content.However, there is a risk of effortlessly generating text using advanced LLMs such as ChatGPT, potentially allowing the completion of class assignments without human effort.As it has a significant impact on the fair evaluation of students, we need to distinguish between text generated by model (model generated text) and written by human (human written text).Detection using existing statistical measures, such as log likelihoods, does not perform well for blackboxed models, such as ChatGPT, because it requires access to the internals of the models.Therefore, we propose a new approach that captures text from two different perspectives using log likelihoods and sentence embeddings with multiple LLMs.In experiments using data, including those generated by the black-box model ChatGPT, our proposed method demonstrated superior accuracy compared to existing approaches.

Read the paper · More papers on PaperTik