The Suffix Tree Construction for Large Sequences over Hand-held Device

Hae‐Won Choi, Hyunsung Kim · 2008

Generally, a suffix tree is an efficient data structure since it reveals the detailed internal structures of given sequences within linear time. However, it is difficult to implement a suffix tree for a large number of sequences in the ubiquitous devices because of the memory constraints. Therefore, in order to compare multi-megabase genomic DNA sequence sets using suffix trees, there is a need to re-construct the suffix tree algorithm. This paper introduces a new algorithm for constructing a suffix tree on the secondary storage with a large number of sequences. Our algorithm divides a suffix tree into three files, in a designated sequence, into parts, storing references to the locations of edges in hash tables.

Read the paper · More papers on PaperTik