Mispronunciation Detection of Mandarin Pronunciation Errors in Tibetan Students Based on BIDFSMN

Zhenye Gan, Wenhao Wei · 2024

This paper is to develop a computer-assisted training system so that mispronunciation by the non-native speaker is recognized and detailed, critical feedback can be given so that second-language learners improve effectively in pronunciation ability. We trained an end-to-end speech recognition model using Bidirectional Deep Feedforward Sequence Memory Networks (BIDFSMN) with the Connectionist Temporal Classification (CTC) method. This method should not rely on phonetic information and forced alignment; instead, it should be used to characterize the extended syllables as the elementary units in detecting pronunciation errors and define 65 types of deviation. Experimental data has shown that this method has relatively high detection accuracy of mistakes and the following indices: 90.74% detection accuracy, 7.53 % rejection rate, and 24.26% acceptance rate. It opens up new views and ways of further developing speech training systems, which have a vast potential for improving the pronunciation of second language learners.

Read the paper · More papers on PaperTik