More Applications of Suffix Trees
Dan Gusfield · Cambridge University Press eBooks · 1997
With the ability to solve lowest common ancestor queries in constant time, suffix trees can be used to solve many additional string problems. Many of those applications move from the domain of exact matching to the domain of inexact , or approximate, matching (matching with some errors permitted). This chapter illustrates that point with several examples. Longest common extension: a bridge to inexact matching The longest common extension problem is solved as a subtask in many classic string algorithms. It is at the heart of all but the last application discussed in this chapter and is central to the k-difference algorithm discussed in Section 12.2. Longest common extension problem Two strings S 1 and S 2 of total length n are first specified in a preprocessing phase. Later, a long sequence of index pairs is specified. For each specified index pair ( i, j ), we must find the length of the longest substring of S 1 starting at position i that matches a substring of S 2 starting at position j . That is, we must find the length of the longest prefix of suffix i of S 1 that matches a prefix of suffix j of S 2 (see Figure 9.1). Of course, any time an index pair is specified, the longest common extension can be found by direct search in time proportional to the length of the match. But the goal is to compute each extension in constant time, independent of the length of the match.