N-gram Indexing for Protein Sequence Databases

Jin-Suk Kim, Mi-Nyeong Hwang · 2009

Motivation: Though the sequence databases of proteins and DNAs are increas- ing in size exponentially, still exhaustive sequence search systems are commonly used in conducting biological researches. However, due to the advancement of information technology, many information retrieval algorithms have been de- veloped to search strings in large-scale text databases and are proved to be successful. We propose that these algorithms could also be applied to the bio- logical data. Results: Four n-gram indexing methods (tri-gram, tetra-gram, penta-gram, and hexa-gram) were applied to extract indices from protein sequences of the PIR-NREF database, and their retrieval effectiveness and speed were mea- sured. Penta-gram method showed the best results that its retrieval effective- ness matches for BLASTP and its retrieval speed was about 38 times faster than BLASTP program. Availability: Our protein sequence search service is accessible at http://proses.kisti.re.kr. Contact: Jinsuk Kim ([email protected])

Read the paper · More papers on PaperTik