Acoustic modeling for native and non-native Mandarin speech recognition
Xin Chen, Jian Cheng · 2012
In this paper, we first described the automatic Spoken Chinese Test (SCT). With a large amount of native and non-native data collected for SCT, different training strategies for acoustic modeling were investigated. Evaluations were performed on native as well as non-native datasets. We discovered that directly combining native and non-native data to train acoustic models did not work well, and the acoustic model trained only on native data achieved better performance when applying to non-native speech. We investigated how to use non-native data effectively, and found that Phonetic Decision Tree (PDT) had a great impact. Discriminative training was found to improve speech recognition accuracy effectively for both native and non-native Mandarin speech.