Vocal and visual design features in multimodal GenAI agents: effects on EFL speaking performance, affect and interaction experience
Chenghao Wang, Xueyun Li, Hui Jin, Bin Zou · Computer Assisted Language Learning · 2026
The integration of multimodal input channels, including text, voice, image interaction, and human-like representations, within Generative AI (GenAI) agents enables immersive speaking experiences by providing linguistic scaffolding and emotional support. Despite their growing use in English as a Foreign Language (EFL) speaking instruction, limited research has examined how vocal and visual design features shape learners’ experiences and outcomes. This study adopts a 2 (Time: pre vs post) × 3 (Group: control, Agent 1, Agent 2) quasi-experimental design to investigate the effects of alternative voice and visual configurations on speaking performance, affective responses, and interactional experience. Two multimodal GenAI agents were developed: Agent 1 employed an animated teacher image with a standard AI-generated voice, whereas Agent 2 featured a realistic teacher photograph with an AI-cloned teacher voice. Participants were drawn from three university EFL classes: the control group (N = 37) received traditional instruction, while experimental group 1 (EG1, N = 34) interacted with Agent 1 and experimental group 2 (EG2, N = 36) interacted with Agent 2. Over eightweeks, learners completed structured speaking tasks under their respective conditions. Multilevel modelling of performance and questionnaire data showed that both EGs outperformed the CG in speaking gains. EG2 demonstrated greater improvements in self-perceived communicative competence and reductions in speaking anxiety, along with higher perceived interactivity and immersion, whereas EG1 showed a significant decrease in foreign language boredom. Follow-up interviews further illustrated how vocal and visual design features shaped learners’ learning experiences. These findings provide insights for the design and pedagogical deployment of multimodal GenAI agents in EFL speaking instruction.