Structured Markov models for speech recognition
Franz Wolfertstetter, G. Ruske · 2002
This paper proposes a new modeling of the structure of speech units as a graph consisting of base functions and a transition network. A cluster algorithm taking into account the actual temporal context of the feature vectors is used to generate the base functions, which are approximated by normal distributions. The subsequent Viterbi-based maximum-likelihood training procedure establishes the transition network and adjusts the transition probabilities. The emerging graphs for the speech units are a structure of branching and recombining trajectory segments describing statistical dependencies in the feature vector sequence within the speech units as well as in the transition regions between them. A speaker-independent evaluation shows the superiority of the proposed modeling compared to mixture-state HMMs, even for an equal number of model parameters.