Scalable top-k retrieval with Sparta
Gali Sheffi, Dmitry Basin, Edward Bortnikov, David Carmel, Idit Keidar · 2020
Many big data processing applications rely on a top-k retrieval building block, which selects (or approximates) the k highest-scoring data items based on an aggregation of features. In web search, for instance, a document's score is the sum of its scores for all query terms. Top-k retrieval is often used to sift through massive data and identify a smaller subset of it for further analysis. Because it filters out the bulk of the data, it often constitutes the main performance bottleneck.