Leveraging Deep Learning for Automatic Pronunciation Assessment in a Mobile Application
Andrei Lesnov, N. Yu. Zueva, Konstantin Turalchuk · 2024
This paper explores the application of the QuartzNet neural network model for pronunciation assessment. We employed the model from the Nvidia NeMo toolkit and fine-tuned it on the LibriSpecch with additional speaker from LJSpeech. To ensure accurate mapping between sounds and linguistic units, we replaced grapheme-based transcriptions with phoneme-based ones. Evaluation on the Speechocean762 dataset revealed a notably high correlation with expert assessments. The paper delves into the specifics and intricacies of the model training and testing process and the stages involved in developing the Android mobile application. The application leverages the pre-trained QuartzNet model and assesses user pronunciation based on the Character Error Rate (CER) between the hypothesis and the reference transcription. QuartzNet's lightweight nature and efficient inference capabilities enabled on-device inference directly on the user's device (Android smartphone). Testing results and comparisons with other pronunciation assessment tools demonstrate the effectiveness and potential of the application for aiding pronunciation improvement and English lanauage learning.