A Selectivity Model for Fragmented Relations: Evaluated for different standard data distributions

Henk Ernst Blok, R.S. Choenni, Henk Ernst, Blok Sunil Choenni, Katarzyna Wac, Henk Blanken, Peter M. G. Apers · Archive ouverte UNIGE (University of Geneva) · 2002

In the estimation of selectivity, many models assume that data is uniformly distributed, which is not true for many applications. In this paper, we discuss a generalized selectivity model, the so-called l##-model which is independent of the data distribution. The model predicts the fraction of a relation that should be selected in order to process a query. We have evaluated this model for di#erent data distributions in order to determine the accuracy of this model. Data distributions that have been considered are the uniform distribution, the normal distribution, the exponential distribution, Pearson's distribution, and Zipf's distribution. From our experiments, it appears that the l##-model predicts the selectivity well, especially for the skewed distributions. Applying the l##-model on di#erent fragment sizes of a relation yields quite acceptable selectivity values as well.

Read the paper · More papers on PaperTik