Leveraging discarded samples for tighter estimation of multiple-set aggregates
Edith Cohen, Haim Y. Kaplan · 2009
Many datasets, including market basket data, text or hypertext documents, and events recorded in different locations or time periods, can be modeled as a collection of sets over a ground set of keys. Common queries over such data, including similarity or association rules are represented as the weight or selectivity of keys that satisfy some selection predicate defined over keys' attributes and memberships in particular sets.