Sequence Profiles and Hidden Markov Models
Birkhäuser-Verlag eBooks · 2006
Profiles are position-specific score matrices derived from multiple sequence alignments. They are typically used for protein classification, where an unknown amino acid sequence is compared to a set of profiles characterizing known protein families. Such comparisons between profiles and anonymous sequences are often more sensitive than pairwise comparisons. Profiles can be rewritten as profile hidden Markov models. Hidden Markov models are based on Markov chains. These consist of a number of states and the probabilities of switching between them. The first order Markov chains considered in this chapter have the property that the chain’s state at some point in time depends only on its direct predecessor state. In hidden Markov models (HMMs), the states of the Markov chain are said to be hidden and to emit observation states, for example a DNA sequence. Given such data and a HMM, there are three classical problems that need to be solved for HMMs to be useful in biological sequence analysis: the scoring, the detection, and the training problem. Efficient solution of these problems leads to many applications of HMMs in biology, including homology detection, for which profiles were originally designed, but also extremely rapid multiple sequence alignment.