Generative AI Evaluation of Human Tutors: Demonstrations and Considerations for Prompt Engineering
Danielle R. Thomas, Jionghao Lin, Sanjit Kakarla, Shambhavi Bhushan, Erin E. Gatz, Shivang Gupta, Ralph Abboud, Kenneth R. Koedinger · 2024
Trained human tutors are in huge demand but short supply, making scenario-based training a valuable tool for providing situational experiences to novices. While tutor training through online lessons has shown significant learning gains, human evaluation of real-life performance is time-consuming and costly. Generative AI, particularly large language models (LLMs), offers a promising solution for assessing complex tutor actions, though research is very limited. This present work explores the practical application of generative AI, focusing on prompt engineering techniques to evaluate human tutors in real-life dialogues. We pull a random selection of 50 transcripts consisting of undergraduate remote tutors working with middle-school students on math. We then assess GPT-4's ability to identify and evaluate tutors providing praise to students and reacting to students making math errors. The use of GPT-4 was found to be proficient in identifying these tutoring situations and assessing tutors application of best practices, highlighting its potential for scalable evaluation. This study provides practical demonstrations of prompt engineering, insights into the use of LLMs for assessing the transfer of learning from training to practice, and discusses the ethical and practical considerations when using LLMs for situational judgment.