SEQUENCE CLASSIFICATION USING HIDDEN MARKOV MODELS

Pranay A Desai · OhioLink ETD Center (Ohio Library and Information Network) · 2005

The field of Bio-Informatics is fast growing with research in various related topics.One such topic is protein sequence classification.This thesis uses this topic as motivation to develop a methodology that uses Hidden Markov Models (HMMs) to classify sequences.Hidden Markov Models are a concept in probability theory widely known for their application in the speech recognition.The three phases of HMMs: training, decoding, and evaluation, are used to classify sequences into clusters that have known similar functional properties.The training phase of HMMs uses a cluster of sequences to learn a model that is most likely to generate the sequences in the training cluster.The decoding and evaluation phases of HMMs use the generated model to calculate the likelihood of an unknown sequence belonging to the same sequence and generating a most-probable path the sequence traverses.The thesis presents background on HMMs along with detailed explanations of the algorithms used to implement all three phases of HMMs.The primary focus of this thesis is on the training phase of HMMs.During the implementation of the training phase we discovered that the phase has a numerical and computational weakness relating to those structures in which some silent states are included as part of the model.The results presented in this thesis test the training algorithms, show their workability and weaknesses, and point towards the silent states related weakness.5.27 Validation File 5 . . . . . . . . . . . . . . . . . .

Read the paper · More papers on PaperTik