Automatic Pronunciation Evaluation and Feedback Generation System based on Resource-efficient factorized TDNN and Phoneme Error Pattern
Il Han, Hyok Kwak, O Chung-Hyok, Kang-Song Choe, Chol-Nam Om · International Journal of Advanced Networking and Applications · 2025
In this paper we describe an automatic pronunciation error detection and feedback generation system for non-native second language learners by using deep acoustic model based on factorized TDNN and language model with phoneme error model. Our system builds language model considering phoneme error patterns and gives a useful feedback for learners. Deep acoustic model consists of TDNN-F with grouped fully-connected layers and shuffle operation. This network architecture maintains recognition accuracy like traditional TDNN and costs less then it. Also, our system evaluates pronunciation proficiency of utterance in word level and phoneme level based on confidence from Minimum Bayesian Risk decoder, feedback is generated on phone error model of L2 learners. This system based on resource-efficient deep acoustic architecture can be deployed in resource-limited mobile devices.