Enhancing GPT-based automated essay scoring: the impact of fine-tuning and linguistic complexity measures

Yingying Liu, Huilei Qi, Xiaofei Lu · Computer Assisted Language Learning · 2025

Recent studies have shown the potential of GPT models in automated essay scoring (AES) and suggested fine-tuning and integrating linguistic complexity measures as key approaches to enhancing model performance. However, research on the reliability, accuracy and fairness of fine-tuned GPT models in AES across different first language (L1) groups of writers is still limited. Additionally, it is unclear whether combining fine-tuning with linguistic complexity measures can further improve model performance. To address these issues, this study evaluates the performance of a fine-tuned GPT model in rating English essays from four L1 groups of writers. We also examine the impact of integrating linguistic complexity measures by comparing three approaches: GPT ratings only, linguistic complexity indices only, and a combination of both. The fine-tuned GPT model achieved 78.3% exact agreement with human raters overall, though demonstrating lower reliability and accuracy for low-level essays, a minority class in the training data. The model showed a general tendency to overrate essays. Its performance varied across the four L1 groups, with the lowest reliability and accuracy observed for L1 German writers. Among the three approaches, GPT ratings outperformed a set of 22 linguistic complexity measures in evaluating essay quality. Integrating GPT ratings with linguistic complexity measures further resulted in marginal improvement. These findings demonstrate the effectiveness of the fine-tuned GPT approach in AES, highlight the importance of addressing L1-related biases, and suggest directions for future research on GPT-based approaches to AES.

Read the paper · More papers on PaperTik