Towards Human-Level Evaluation: Assessing the Potential of GPT-4 in Automated Evaluation and Feedback Generation on Japanese Essays

Sayaka Nakamoto, Yoshihiro Okamoto, Takashi Nakakouchi, Kazutaka Shimada · 2024

In recent years, Automated Writing Evaluation (AWE) has been extensively researched within the field of AI in education. This paper explores generative AI, such as GPT-4, which has garnered significant attention for its ability to score essays and provide feedback to students. We designed prompts for GPT-4 to assign scores and rationales based on a given rubric and to generate feedback beneficial for students' development. We compared the evaluations produced by GPT-4 with those made by human evaluators. The results demonstrate GPT-4's potential to assist in generating evaluations at a human level. In addition, we analysed the consistency of the scoring and the quality of the rationales and feedback generated by GPT-4. In this paper, we will share our analysis and also describe the points that need to be improved for implementation in practice.

Read the paper · More papers on PaperTik