AI-Based English Vowel Formant Recognition and Pronunciation Correction Platform
Jianying Deng, Chuang Liu, Zixin Zhou · 2024
English vowel pronunciation is very challenging for non-native English learners, as it requires precise articulation often guided by subjective feedback without objective standard. Existing tools give overall scores based on speech recognition, but lack real-time corrective feedback for articulatory details like tongue position, placement, and lip shape. They also show that there is no well-established digital standard for the measurement of English vowel pronunciation, which can hinder reliability in training for this reason. In this study, we present a pronunciation correction platform based on AI and formant analysis (F1, F2, F3) for providing real-time feedback and visualization of the articulatory adjustments. Over the course of their study, learners get feedback that gives them precise direction on tongue height, front-back location, as well as degree of lip rounding, to conform more closely to standard pronunciation. It introduces three principal innovations: (1) Fine-Grained Formant Analysis, which provides informative feedback that connects different articulatory features to very specific linguistic patterns; (2) Visualization Feedback, modeling in a visual interface the detailed deviations from target pronunciation; (3) AI-Driven Personalization, offering automated personalized feedback generated dynamically based on an individual's learning needs. This platform fills a significant gap by creating a digital standard for vowel pronunciation and enhancing pedagogy through detailed feedback as well as visualization. Its applications include, but are not limited to, language learning, speech therapy, and accent modification.