Investigating the Potential of Large Language Models for Automated Writing Scoring
Shan Wang · Atlantis Highlights in Computer Sciences/Atlantis highlights in computer sciences · 2024
This study investigates the potential of large language models (LLMs), specifically GPT-4, for automated writing scoring and feedback generation.Employing a mixed-methods approach, the research evaluates the accuracy and reliability of GPT-4 in predicting essay scores and the quality of its generated feedback.The results demonstrate a high level of agreement between GPT-4 scores and human raters, as evidenced by the confusion matrix and Quadratic Weighted Kappa metric.Qualitative analysis of GPT-4 feedback suggests its ability to provide constructive and comprehensive suggestions for improving student writing.However, there are still limitations surrounding LLM-based automated scoring and feedbacks.Thus, this study proposes the use of LLM-based systems as formative assessment tools to complement human judgment.