Leveraging Machine Learning for Automatic Pronunciation Assessment in English Language Learning
C. Chamundeshwari, S Premalatha, Yaisna Rajkumari, Marwan Aziz Mohammed, T. Sathyaseelan, M Vimochana. · 2025
Automatic Pronunciation Assessment (APA) in English Language Learning (ELL) has been observed to be indispensable for developing the learners’ talking fluency, as well as helping them obtain feedback which will enable them enhance their pronunciation patterns appropriately. Automatic assessment of pronunciation is highly advantageous for non-native learners because it addresses the challenges implied by the process and supports the learning process by pointing out the exact mistakes made by the learner. However, prior studies have had some problems in attaining high recognition accuracy especially in pronunciation evaluation since conventional methods as HMM and basic DNN cannot adequately model temporal dynamics in speech data intensity. They also tend not to operate sub optimally in noisy conditions or where the speaker’s accents are different. In order to overcome these limitations, the present study proposes a new method of automatic pronunciation assessment based on Convolutional Recurrent Neural Networks (CRNNs). While convolutional layers extract the visual features to represent the input image, the recurrent layers successfully continue utilizing the temporal features of the raw speech for more precise pronunciation assessment as compared to the traditional CRNN model. The proposed method is implemented in Python and showed excellent performance, as it was trained on the large dataset; the accuracy is $\mathbf{9 9. 4 \%}$. As the outcomes shown, employing CRNN-based model has an improvement compared to the conventional approach in terms of both accuracy and scalability in the pronunciation assessments of English language learning.