Distributed Non-Parametric Representations for Vital Filtering: UW at TREC KBA 2014

Ignacio Cano, Sameer Kumar Singh, Carlos Guestrin · 2014

Identifying documents that contain timely and vi-tal information for an entity of interest, a task known as vital filtering, has become increasingly important with the availability of large document collections. To efficiently filter such large text corpora in a streaming manner, we need to com-pactly represent previously observed entity con-texts, and quickly estimate whether a new doc-ument contains novel information. Existing ap-proaches to modeling contexts, such as bag of words, latent semantic indexing, and topic mod-els, are limited in several respects: they are un-able to handle streaming data, do not model the underlying topic of each document, suffer from lexical sparsity, and/or do not accurately estimate temporal vitalness. In this paper, we introduce a word embedding-based non-parametric repre-sentation of entities that addresses the above limi-tations. The word embeddings provide accurate and compact summaries of observed entity con-texts, further described by topic clusters that are estimated in a non-parametric manner. Addition-ally, we associate a staleness measure with each entity and topic cluster, dynamically estimating their temporal relevance. This approach of using word embeddings, non-parametric clustering, and staleness provides an efficient yet appropriate rep-resentation of entity contexts for the streaming setting, enabling accurate vital filtering. 1.

Read the paper · More papers on PaperTik