Automatic conversational assessment using a GPT-based model for EAP speaking

Vahid Aryadoust, Joann Wong, Wenxin Zhang, Jiang Qian, Sai Mei Zhang, Tingting Liu, Yining Han · English for Specific Purposes · 2026

This study examined how a conversational agent, customized through advanced prompting in OpenAI’s GPT Builder, can be used for speaking assessment. The system was developed around four key dimensions: measurement definition, system design, performance evaluation, and learner experience. First-year English majors from two large universities participated in assessments where students could choose topics, control response time, and request clarification. Using many-facet Rasch measurement and linear mixed-effects models, we found that GPT assigned slightly lower scores than human raters overall, though mean score differences were not statistically significant. Three of eight scoring categories showed statistically significant differences between GPT and human ratings. Nonetheless, GPT’s scoring demonstrated acceptable consistency and reliability and exhibited typical rater tendencies such as restrained use of extreme scores. Students requesting clarification received lower scores, while those taking longer performed slightly better. Students’ perceptions about GPT did not significantly influence scores, which suggests that the test outcomes were not influenced by factors unrelated to speaking ability. Overall, findings indicate that GPT-powered systems may provide a useful and scalable option for formative speaking assessment, especially in classroom settings prioritizing student choice and flexibility.

Read the paper · More papers on PaperTik