Development and validation of an AI-enhanced multimodal training program: evidence from non-native Mandarin tone learning
Bing Cheng, Kangzhi Liao, Yilin Xiang, Yu Zou, Xiaojuan Zhang, Yang Zhang · Computer Assisted Language Learning · 2025
Computer-assisted language learning (CALL) technologies offer unprecedented opportunities to tackle challenges in second language (L2) phonological acquisition. This study introduces an artificial intelligence-enhanced multimodal production training program for Mandarin tone learning that addresses critical gaps in existing technologies by integrating automatic speech recognition and dynamic time warping algorithms to provide learners with immediate, objective pronunciation analysis and personalized visual feedback. To assess the effectiveness of the training program, we analyzed the production and perception skills among non-tonal L2 learners in pre- and post-tests. Detailed acoustic analyses of monosyllabic and disyllabic words showed that participants in the training group demonstrated significantly closer alignment with native speakers’ pronunciations for both trained and untrained word types, while perceptual assessments revealed improved word identification accuracy. Our findings demonstrated the transformative potential of theory-grounded, technology-driven multimodal training systems in advancing L2 phonological skills, providing a promising framework for future CALL interventions that prioritize both pedagogical effectiveness and learner accessibility.