Sequence Homology Search Based on Database Indexing Using the Profile Hidden Markov Model
Qiang Xue, James Cole, Sakti K. Pramanik · 2006
The profile hidden Markov model (PHMM) has received increasing attention in the field of protein homology detection, since profile-based methods are much more sensitive in detecting distant homologous relationships than pairwise methods. Pure dynamic-programming-based systems are often used for PHMM searches. However, these dynamic-programming- based systems are very time consuming for a large database. For instance, it may take approximately 15 minutes to search a short model of length 12 in the GenBank protein sequence database. Instead of searching the database sequentially, we search the database based on a tree-structured database indexing, called the HD-tree. The HD-tree is able to reduce the PHMM search time significantly without reducing the quality of search results. Performance of search using the HD-tree is compared with that of HMMER, a popular implementation of PHMM for protein sequence analysis. It is shown that the HD-tree approach is orders of magnitude faster than HMMER for short queries