The Effectiveness of Using Large Language Model to Generate English Reading Comprehension Diagnostic Test: Generative AI and Human-Made comparison

Pimvaree Khamrassamee, Putcharee Junpeng, Thanapong Intharah · 2025

This study compares an AI-generated English Reading Comprehension test (Claude) with a human-created version on the topic of places. The research aims to evaluate the quality of AI-generated assessments with human design. Data from 128 ninth-grade students were analyzed using the Rasch model. Findings indicate that the AI-generated test demonstrated comparable content validity (CVI ratings consistently 3–4) to the human-made version while exhibiting superior reliability within acceptable psychometric limits. In a Turing Test, two English proficiency experts were more successful in identifying AI-generated items compared to experts from other fields. These results support the potential of AI in developing valid and extensible educational assessments.

Read the paper · More papers on PaperTik