An Early Review of Generative Language Models in Automated Writing Evaluation: Advancements, Challenges, and Future Directions for Automated Essay Scoring and Feedback Generation
Huang Yue, Corey Palermo, Ruitao Liu, Yong He · Chinese/English Journal of Educational Measurement and Evaluation · 2025
Automated writing evaluation (AWE) has long supported assessment and instruction, yet existing systems struggle to capture deeper rhetorical and pedagogical aspects of student writing. Recent advances in generative language models (GLMs) such as GPT and Llama present new opportunities, but their effectiveness remains uncertain. This review synthesizes 29 studies on automated essay scoring and 14 on automated writing feedback generation, examining how GLMs are applied through prompting, fine-tuning, and adaptation. Findings show GLMs can approximate human scoring and deliver richer, rubric-aligned feedback, but fairness, validity, and ethical issues remain largely unaddressed. We conclude that GLMs hold promise to enhance AWE, provided that future work establishes robust evaluation frameworks and safeguards to ensure responsible, equitable use.