Speaker dynamics as a source of pronunciation variability for continuous speech recognition models

Rebecca Bates · 2004

A significant source of variation in spontaneous speech is due to intra-speaker pronunciation changes. Previous work has identified several factors related to pronunciation variability, such as phonetic context and speaking rate, which are useful to model in automatic speech recognition. This work examines new higher-level information sources: syntax, discourse structure and prosody, specifically the relationship between these factors and pronunciation variation as seen in reduction and hyper-articulation. The key contributions of this work include 1) analysis of high-level factors, providing new cues for improving prediction of pronunciation variation, 2) a framework for including dynamic pronunciation models in automatic speech recognition systems, and 3) an analysis of feature-based pronunciation models with suggestions for their incorporation into ASR systems. Key findings from the analysis of high-level factors are attributes that are most useful for predicting variability, including: part-of-speech (POS) of the target word and neighboring words, location of the word in an utterance, the number of F0 slope changes within the word, word duration, and average word energy. Pronunciation prediction experiments show a reduction in phone error rate of 2.3% relative and similar reductions in perplexity over a baseline model using only phonetic context. Incorporating higher-level information (such as hypothesis-dependent word context or word-level F0 values) into ASR systems requires a rescoring approach. A framework for this is presented, with recognition results using various types of pronunciation models on the Switchboard task. We obtain a small but statistically significant...

Read the paper · More papers on PaperTik