Examining the Voice-Image Matching for Pedagogical Agents Presented in Instructional Videos
Bin Jing, Changcheng Wu, Zhongling Pi, Yu Zhou, Hongliang Ma · The Journal of Experimental Education · 2025
Technological advancements have made on-screen pedagogical agents (PAs) increasingly easy to implement. However, no consistent conclusion exists regarding how these PAs are presented in instructional videos, particularly concerning matching their voices and images. This study examined the influence of the consistency effect of PA’s voice-image matching on learners’ outcomes in video learning. We first created a set of four robot images for a pretest and recruited 163 college students to select their preferred robot image. The most selected image was chosen as the visual appearance of the robot PA for the subsequent experiment. We then conducted a 2 (images: human vs. robot) × 2 (voices: human vs. synthesized) between-subjects experiment to collect eye movement data from 110 college students. Learning outcomes were evaluated based on social presence, persona perception of PA, and learning performance. This study indicated that matching the human voice to the human image improved learners’ attention allocation on learning content, social presence, human-like perception of PA, and learning performance. However, no positive effects were revealed when the modern computer-synthesized voice was matched with the robot image. Therefore, we recommend instructors prioritize using human voices and images when designing PAs in instructional videos.