KNITTIR: Syntactical Text Indexing for Analytics

Thanh-Hi Anthony Vu, Dhruv Gupta · 2024

Scalable text analytics requires retrieval of similar text regions spread across millions of documents. It also requires that we can reason about entities by categorizing them; contrasting them to other entities; and ranking them using time and numbers. We present KNITTIR that assists in such complex text analytical tasks. KNITTIR uses semantic annotations such as parts-of-speech, named entities, and their syntactical relationships to words in text. To simplify text analytics, KNITTIR uses a new search framework wherein a vertical partitioning of semantically annotated text is used. This allows users to aggregate and manipulate evidences to their queries from multiple text regions spread across millions of documents. To scale analytical queries, KNITTIR creates indexes using the syntactical modeling of annotated text. Our experiments over 22 million documents show that KNITTIR obtains speedups of up to 60× for performing similarity search and up to 69K× for reasoning tasks.

Read the paper · More papers on PaperTik