University of Texas School of Information at TREC 2007.

Miles Efron, Don Turnbull, Carlos Ovalle · Text REtrieval Conference · 2007

Our system used Apache’s Lucene library [1] for its core indexing and retrieval functions. We also relied on language modeling extensions to Lucene provided by the Informatics Institute at the University of Amsterdam [2]. However, We altered these libraries to enable our IR approach. In particular, our results rely on a variant of the Kullback-Leibler (KL) divergence model [3, 4]. Given a query q we derive a score for each feed f in the corpus by the negative KL-divergence between the query language model and the language model for f. In the interest of maximizing precision at low numbers of documents retrieved, we limited our analysis to each feed’s RSS posts, as opposed to its complete HTML representation.

Read the paper · More papers on PaperTik