Evaluating the Efficacy of Large Language Models in Automating Academic Peer Reviews

WeiminZhao, Qusay H. Mahmoud · 2024

This paper explores the application of large language models (LLMs) in automating the peer review process for academic papers, a critical area for enhancing the efficiency and consistency of scholarly publication. We utilized GPT-4-0125 to automatically generate reviews for 20 papers sourced from openreview.net and analyzed the AI-generated peer reviews for quality and effectiveness. The analysis includes a detailed assessment of the text properties of the reviews, such as sentiment, revealing that LLM-generated reviews tend to be more uniformly positive than their human-written counterparts. In addition, we conducted a user survey in which participants attempted to distinguish between AI-generated and human-written reviews. The survey results indicated a low correct identification rate, suggesting that participants often could not discern the origin of the review, thereby highlighting the potential of LLMs to mimic human-like review qualities. However, the study also identifies limitations in LLM's performance, particularly concerning the variability in review quality, which appears to correlate with the model's vocabulary usage of the generated content.

Read the paper · More papers on PaperTik