Human Evaluation of Large Language Models: A Review and Protocol Selection Framework
Tad T. Brunyé · AI · 2026
Evaluating large language models (LLMs) critically depends on human judgment. This article reviews and develops a conceptual framework for human-centered LLM evaluation, synthesizing research across evaluation methodology, psychometrics, cognitive science, and domain-specific applications. Four primary challenges are identified that limit current human evaluation practice: imperfect gold standards, evaluator fatigue and overload, shared and unique bias structures across humans and LLM judges, and the routine omission of uncertainty and dispersion estimates. To address these gaps, the STEP-V design framework is proposed: Stakes, Task-type, Evaluator availability, Purpose, and Volume, for selecting human and/or automated LLM evaluation methods under real-world constraints. An evaluator failure mode taxonomy is also proposed that analyzes human and LLM judges within a common error framework, clarifying where hybrid pipelines can compensate for weaknesses and where they might compound them. The framework motivates a more rigorous science of LLM evaluation, one that treats human judgment as a necessary but fallible measurement requiring explicit design, calibration, and uncertainty quantification.