Detecting Reference Errors in Scientific Literature with Large Language Models

Tianmai M. Zhang, Neil F. Abernethy · PubMed · 2024

Reference errors, such as citation and quotation errors, are common in scientific papers. Such errors can result in the propagation of inaccurate information, but are difficult and time-consuming to detect, posing a significant threat to the integrity of scientific literature. To support automatic detection of reference errors, we evaluated the ability of large language models in OpenAI's GPT family to detect quotation errors. Specifically, we prepared an expert-annotated, general-domain dataset of statement-reference pairs from journal articles, one-third of which is in biomedicine. Large language models were evaluated in different settings with varying amounts of reference information provided by retrieval augmentation. Results showed that large language models are able to detect erroneous citations with limited context and without fine-tuning. This study contributes to the growing literature that seeks to utilize artificial intelligence to assist in the writing, reviewing, and publishing of scientific papers as well as grounding of language model responses.

Read the paper · More papers on PaperTik