On automated query modification techniques for databases
Kaizheng Du · OhioLink ETD Center (Ohio Library and Information Network) · 1993
In many cases, in order to satisfy certain constraints specified in a database query, the query needs to be modified. However, users' modification is an extra burden on users and, sometimes, may not be correct when users lack knowledge about the database. In a real-time environment, users' manual modification may be too slow to make a query satisfy a given time dead-line. In this thesis, a general automated database query modification model is proposed for automatically modifying database queries with constraints. Five types of query modification constraints, namely, time constraints, error constraints, aggregate function constraints, PCF constraints and count proportionality constraints are introduced. To enable query modification to be performed automatically, two query modification protocols, namely, the use of superset/subset chains based on relation fragmentation and the use of sampling data, are specified. Based on the two query modification protocols, a detailed query modification mechanism for enforcing time and error constraints are given. An iterative query evaluation technique is used to process nonperiodically occurring queries with time constraints. An incremental query evaluation technique is used to process periodically occurring queries with time constraints. Error removal and error estimation techniques are introduced to enforce error constraints. Finally, extending the above listed techniques into enforcing the remaining three types of constraints is briefly summarized. In this thesis, query estimation techniques for aggregate relational algebra queries COUNT, SUM and AVERAGE are also presented. Two statistical estimators, the Jackknife estimator and the Chao's estimator, for COUNT queries with projection are used. Estimators using double sampling technique for SUM and AVERAGE queries are introduced, and new sampling plans based on systematic sampling and stratified sampling to increase the estimation accuracy are investigated. These new estimators and sampling plans are extensively used in automated query modification. Some of the techniques and associated algorithms proposed in this thesis have been implemented in CASE-DB, which is a prototype relational database management system developed at Case Western Reserve University.