Reflective Prompt Engineering for Assessment Rubric Optimization: An Empirical Study of Human–AI Alignment
Edgaras Norgaila, Linda Daniela, Daiga Kalniņa · Technology Knowledge and Learning · 2026
The rapid integration of GenAI into education has created opportunities and challenges for student assessment. This study explores the application of Reflective Prompt Engineering (RPE), a novel and innovative framework designed to enhance automated essay scoring through iterative rubric refinement, model self-analysis, and reflective justification of scoring decisions. Unlike conventional prompting methods, the RPE framework is designed for iterative refinement, where the model’s scoring decisions are reflected upon, critiqued, and adjusted to achieve closer alignment with human evaluative reasoning. Using GPT-4o and GPT-5-chat, 24 student essays were evaluated against rubrics designed by instructors and refined through the automated RPE process. Findings demonstrate that GPT-4o exhibited closer alignment with human raters in first-pass scoring, while GPT-5-chat initially inflated scores but achieved partial bias reduction after refinement. However, one cycle of rubric refinement and self-analysis did not consistently improve accuracy, introducing instability and reduced agreement. These findings suggest that while RPE offers a structured approach to integrating reflection into scoring, its effectiveness remains limited in practice when evaluating essays. Future research should build on larger datasets to examine whether RPE can more effectively capture multidimensional essay qualities such as argumentation, originality, and critical thinking.