An Automatic Phoneme Recognizer for Children’s Filipino Read Speech

Francis D. Dimzon, Ronald M. Pascual · 2020

In this paper, we developed a phoneme-level speech recognizer using hidden Markov models (HMMs) for children's Filipino read speech. Forced alignment was done resulting to 60% of phoneme locations were within 20ms threshold agreement with human transcriptions. Experiments showed that the highest accuracy of 57.47% was attained using a 5-active state HMM model. Increasing the number of Gaussian mixtures from 1 to 6 gave an increased accuracy of 10%. There was also an increase of 2% accuracy when vocal track length normalization was implemented. Further research may involve improving the accuracy of detection by increasing the number HMM's active states, varying the number of Gaussian mixtures for the training data set, integrating a language model, and implementing hybrid Deep Neutral Network (DNN)-HMM models. Moving from phoneme- to word-level children's Filipino speech recognizer provides a more comprehensive way to design an automatic children's oral reading fluency assessment system.

Read the paper · More papers on PaperTik