Comparative Analysis of GPT and BERT for Automated Open-Ended Question Scoring

Sherlyn Hwang, Saaveethya Sivakumar, Choo W. R. Chiong · 2025

Automated scoring systems based on machine learning are well-established, significantly reducing educators' grading workload. Unlike closed-ended questions, scoring open-ended questions requires analysing and evaluating students' answers, which is time-consuming and tedious. Especially in having a large student population, educators cannot give timely feedback. A reliable automated scoring of open-ended question systems can reduce the workload and consistency of the grading process, allowing them to focus more on teaching and supporting students. The automated scoring system is built in machine learning models, learning human decisions from the provided dataset. The state-of-the-art models in automated grading involve the GPT and BERT models, performing well on NLP tasks. This paper benchmarked the performance of the GPT-3.5 model against existing BERT models in scoring short answer questions using Mohler dataset.

Read the paper · More papers on PaperTik