Algorithms on Stings, Trees, and Sequences
Dan Gusfield · ACM SIGACT News · 1997
The history and motivationAlthough I didn't know it at the time, I began writing this book in the summer of 1988 when I was part of a computer science research group at the Human Genome Center of Lawrence Berkeley Laboratory.Our group followed the standard assumption that biologically meaningful results could come from considering DNA as a one-dimensional character string, abstracting away the reality of DNA as a flexible three-dimensional molecule, interacting in a dynamic environment with protein and RNA, and repeating a life-cycle in which even the classic linear chromosome exists for only a fraction of the time.A similar, but stronger, assumption existed for protein, holding for example that all the information needed for correct three-dimensional folding is contained in the protein sequence itself, essentiaUy independent of the biological environment the protein lives in.This assumption has recently been modified, but remains largely intact.For non-biologlsts, these two assllmptions were (and remain) a god-send allowing rapid entry into an exciting and important field.Statements such as "The digital information that underlies biochemistry, cell biology, and development can be represented by a simple string of G's, A's, T's and COs.This string is the root data structure of an organism's biology."reinforced the importance of sequence-level investigation.So without worrying much about the more diIBcult chemical and biological aspects of DNA and protein, our computer science group was empowered to consider a variety of biologically important problems defined plrimarily on sequences, or (more in the computer science vernacular) on strings.We organized our efforts into two high-level tasks.First, to learn the relevant biology, laboratory protocols, and existing algorithmic methods used by biologists.Second to canvass the computer science literature for ideas and algorithms that weren't already used by biologists, but wkich might plausibly be of use either in current problems, or in problems that we could anticipate arising when vast quantities of sequenced DNA or protein become available.