Enhancing Short Answer Grading with OpenAI APIs
Sebastian Speiser, Annegret Weng · 2024
Automated short-answer grading can accelerate and standardize the assessment of tests in higher education. The topic has received a significant boost due to the rapid development of powerful LLM models in recent years. We examine the performance of the OpenAI models GPT-3.5 and GPT-4o on the CSSAG dataset and demonstrate that GPT-4o, in particular, achieves an accuracy that falls within the range observed in assessments by different human evaluators. We pay special attention to cases where there are significant deviations from the reference assessment. Additionally, we discuss the practical implications.