TopX: Efficient Top-k Query Processing for Text, Semistructured, and Structured Data

Martin Theobald, Gerhard Weikum, Norbert Fuhr · MPG.PuRe (Max Planck Society) · 2005

TopX is a top-$k$ retrieval engine for text and XML data. Unlike Boolean engines, it stops query processing as soon as it can safely determine the $k$ top-ranked result objects according to a monotonous score aggregation function with respect to a multidimensional query. The main contributions of the thesis unfold into four main points, confirmed by previous publications at international conferences or workshops: \begin{itemize} \item Top-$k$ query processing with probabilistic guarantees. \item Index-access optimized top-$k$ query processing. \item Dynamic and self-tuning, incremental query expansion for top-$k$ query processing. \item Efficient support for ranked XML retrieval and full-text search. \end{itemize} Our experiments demonstrate the viability and improved efficiency of our approach compared to existing related work for a broad variety of retrieval scenarios.

Read the paper · More papers on PaperTik