Modeling acoustic-phonetic detail in an HMM-based large vocabulary speech recognizer

Li Deng, Matthew Lennig, V. Gupta, Paul G. Mermelstein · 2003

The acoustic recognizer of the INRS-Telecommunications 60000-word-vocabulary isolated-word recognition system is discussed. The task of the acoustic recognizer is to generate a list of word hypotheses and their likelihoods based on the acoustic data for each input word. Two sets of experiments are reported in which such knowledge is incorporated into the hidden Markov models (HMMs) used during recognition. In the first set, vowel duration properties are used in the HMMs. In the second set, word-initial and word-final stop consonants are modeled as a sequence of context-dependent subphonemes. The performance of the recognizer is significantly improved by appropriate utilization of the context-dependent vowel-duration information and the context-dependent microsegmental properties of stop consonants. >

Read the paper · More papers on PaperTik