Estimating selectivities in data bases
Stavros Christodoulakis · 1982
In this thesis we examine the problem of modelling data base contents and data placement on devices. This modelling is necessary in analytic data base performance evaluation studies in order to estimate the number of records of a file that have to be retrieved in response to the user(s) requests, as well as the number of blocks of the file containing these records. The cpu, io, and telecommunication costs of the system are directly or indirectly expressed in terms of these quantities. We first show that certain assumptions used for modelling data base contents, data placement on devices and user requests often are not satisfied in actual data base environments, and that they may lead to errors in model predictions. We examine formally implications of non-uniformity and dependencies of attribute values in data base design and data base performance evaluation. Thereafter, we provide more detailed modelling techniques based on multivariate statistical model, and we demonstrate their use in improving data base performance.