Acoustic and language modeling for children's read speech assessment
Hitesh Tulsiani, Prakhar Swarup, Preeti Rao · 2017
Automatic speech recognition can be used to evaluate the accuracy of read speech and thus serve a valuable role in literacy development by providing the needed feedback on reading skills in the absence of qualified teachers. Given the known limitations of ASR in the face of insufficient task-specific training data, the selection of acoustic and language modeling strategies can play a crucial role in achieving acceptable performance in the task. We consider the problem of detecting mispronunciations in read-aloud stories in English (as a second language) by children in the age group 10-14. Multiple available datasets that separately capture the characteristics of children's speech and Indian accented English are used to train and adapt the acoustic models. A knowledge of the text together with the prediction of mispronunciation errors helps to define an effective language model. We present mispronunciation detection performance on a small test dataset of field recordings and discuss implications for further work.