Real-Time Pronunciation and Language Learning Assistance using Deep Learning
Gontla Venkata Monesh, Gadi Joshith, Susanna M Santhosh, V Harish, Suja Palaniswamy · 2025
This work provides an interactive platform for pronunciation evaluation using deep learning techniques. Here, the text datasets are converted into audio, and then users can input their own audio recordings of pronunciations that are compared with reference pronunciations text data. The real-time transcription, phoneme recognition, and accent detection models used here are Wav2Vec 2.0, DeepSpeech, and LSTM hybrids. Methods of feature extraction that enhance the accurate pronunciation grading include MFCCs and spectrograms. Key features include phonetic analysis, real-time audio-visual feedback, and gamified exercises to improve learner engagement. With this emphasis on phoneme accuracy, the system provides very granular views of pronunciation patterns and how to tackle specific areas for improvement.