Genetic sequence data retrieval and manipulation based on generalized suffix trees

Paul Bieganski · 1995

The volume of genetic sequence data available to scientists is rapidly increasing. Sequence information is most commonly stored in computer memory in contiguous locations, in order of the molecules in the biological sequence. This storage method is not efficient for many sequence information processing applications. Generalized Suffix Trees may provide a system capable of content-based addressing of sequence. Such a system allows a sequence to be accessed in terms of what it contains without having to specify where it is contained. Our thesis is an investigation of the applicability of Generalized Suffix Trees (GSTs) to content-domain processing of genetic sequence data. We explore the algorithms involved in using GSTs for sequence data processing and their implementation. We identify the areas of applicability of GSTs to genetic sequence processing. We define a GST algebra providing the operations necessary for describing a large class of sequence processing applications. We introduce a new algorithm for sequence homology searching based on GST alignment. We perform theoretical analysis of GST-based sequence processing and search algorithms and verify the results through experiments. We review GST implementation issues and develop a GST sequence analysis toolkit. Finally, we implement a number of GST-based applications and evaluate their performance in selected research problems.

Read the paper · More papers on PaperTik