Bayesian computation for hidden Markov models
Lindsay Anne Foreman · Spiral (Imperial College London) · 1994
In order to motivate the use of hidden Markov models (HMMs) for classification tasks, a general background to the field of automatic speech recognition (ASR) is provided and used as a point of reference throughout the thesis.The search for improvements to classical template-matching procedures in ASR has led to the adoption of the HMM, which offers a more formal and flexible statistical framework within which to conduct pattern recognition.The HMM is represented by an observable process, associated with which is an underlying unobserved Markov chain state process.A detailed description of the HMM is given along with the problems to be addressed in its implementation, such as parameter estimation, classification procedures and recovery of the unobserved Markov chain state process.Current statistical methods for dealing with these problems are reviewed.These include a version of the EM algorithm used for maximum likelihood parameter estimation, a search procedure called the Viterbi algorithm, which uses dynamic programming techniques to recover the optimal unobserved state sequence underlying the observed data, and methods of comparing observed patterns and HMMs to find the best match for classification purposes.The Viterbi algorithm is then studied in more depth and a generalisation of the algorithm is established.Examples are given to illustrate its usefulness, particularly for classification purposes in handwritten text recognition and in a pharmaceutical application.The concepts underlying the Bayesian philosophy are introduced; in particular, a Bayesian approach to parameter estimation in HMMs.Markov chain Monte Carlo (MCMC) stochastic simulation techniques are discussed, including the Gibbs sampler and Hastings algorithm, which are useful computational tools in the Bayesian analysis of HMMs.Approaches to the design of efficient MCMC samplers are then critically reviewed.These MCMC methods are then implemented and various sampling designs investigated in a Bayesian approach to HMM parameter analysis.B Appendix to Chapter 3 192 B.l Model parameters for the text recognition example B.2 Model parameters for the EEG data example C Appendix to Chapter 5 C.l State sequence labels C.2 Simulation of ordered Normal variables References 205 signal itself to label the word boundaries.Variability Speech signals exhibit various forms of variability which can be divided into two types: intra-speaker and inter-speaker.Inter-speaker variability is that which is attributable to differences between speakers, due to factors such as sex, age and local accent.However, restricting consideration to a single speaker does not eliminate all variability.Such intra-speaker variability comes in the form of differences in the volume or pitch at which words are spoken, or the health and mood of the speaker.Also, the context in which words are spoken governs which parts are accentuated and the length of time spent on pronouncing them.In fact, a single speaker pronouncing the same sentence on different occasions will never produce identical speech signals. ContextBasing recognition solely upon vocal sounds throws up problems in differentiating between words which sound similar, e.g."to", "too" and "two".Thus, recognition must be performed in context -that of a sentence, say, under grammatical constraints -taking into account the broader meaning behind the communication (semantics) of which speech is only one part.The development of adequate data-based recognisers therefore rests on our ability to overcome such problems. Template-matching recognition techniques