Parallel Suffix Arrays for Corpus Exploration

Johannes Goller · Institutional Repositories DataBase (IRDB) · 2010

This paper describes how recently developed techniques for suffix array construction and compression can be expanded to bring a new data structure, called parallel suffix array, into existence, which is suitable as an in-memory representation of large annotated corpora, enabling complex queries and fast extractions of the context of matching substrings.It is also shown how parallel suffix arrays are superior to existing corpus search engines, in particular when sequential queries and corpora that are hard to tokenize are involved.

Read the paper · More papers on PaperTik