A Two-Phase MapReduce Algorithm for Scalable Preference Queries over High-Dimensional Data

Gheorghi Guzun, Guadalupe M. Canahuate, David Chiu · 2016

Preference (top-k) queries play a key role in modern data analytics tasks. Top-k techniques rely on ranking functions in order to determine an overall score for each of the objects across all the relevant attributes being examined. This ranking function is provided by the user at query time, or generated for a particular user by a personalized search engine which prevents the pre-computation of the global scores. Executing this type of queries is particularly challenging for high-dimensional data. Recently, bit-sliced indices (BSI) were proposed to answer these high-dimensional preference queries efficiently in a centralized environment.

Read the paper · More papers on PaperTik