Applying psychometric methods to distinguish between human and generative AI responses to multiple-choice assessments
Alona Strugatski, Giora Alexandron · Computers and Education Artificial Intelligence · 2026
The growing use of generative AI (GenAI) tools like ChatGPT raises serious concerns about academic integrity, especially in the context of assessments. While detection efforts have focused on open-ended responses, multiple-choice questions (MCQs), which are common in high-stakes testing, remain largely overlooked, partly due to their perceived detection difficulty. The present work establishes a theoretical and empirical foundation for applications of psychometric theory to separate GenAI and human responses. Specifically, it investigates whether person-fit statistics (PFS), a class of methods within Item Response Theory (IRT) used to evaluate how well an examinee’s response pattern fits the expectations of the IRT model, can distinguish between GenAI and human responses to MCQ assessments. To study this, we use data from two authentic assessment contexts: a high-school level chemistry test and the national university entrance exam, each with approximately 1000 human respondents. Our results demonstrate that PFS reveal significant differences between human responses and those generated by advanced chatbots (ChatGPT, Claude, and Gemini), which appear as ‘aberrant’ respondees. We also demonstrate that different chatbots present significantly different response patterns, suggesting that they should be treated as a heterogeneous group of ‘intelligences’ rather than as a single one. Using the PFS measures, we also demonstrate, however, that newer GenAI versions not only improve in performance but also become more ‘human-like’ in their response patterns. Together, these results position IRT as a robust framework for characterizing and separating human and GenAI response patterns in MCQ assessments, providing a theoretical foundation and empirical evidence.