Hierarchical methods in automatic pronunciation evaluation

Joseph Tepperman · University of Southern California Digital Library · 2015

Technology that can automatically categorize pronunciations and estimate scores of pronunciation quality has many potential applications, most notably for second-language learners interested in practicing their pronunciation along with a machine tutor, or for automating the standard assessments elementary school teachers use to measure a child's emerging reading skills. The many sources of variability in speech and the subjective perception of pronunciation make this a complex problem. Linguistic hierarchies - in speech production, perception, and prosodic sturcture -- help to conceive of the variability as existing on multiple simultaneous scales of representation, and offer an explanatory order of precedence to those scales. These theories are beginning to gain widespread attention and use in traditional speech recognition, but experimenters in pronunciation evaluation have been slow to embrace them. This work proposes using theories of hierarchical structure in speech to inform a chosen computational framework and scale of analysis when performing automatic pronunciation evaluation, on the assumption that they will offer improvements over non-hierarchical methods and can be used to rate pronunciation with performance comparable to that of inter-human agreement.

Read the paper · More papers on PaperTik