Automatic generation of type abstraction hierarchies for cooperative query answering
Kuorong Chiang · 1995
A cooperative query answering (CQA) system is a database system that provides approximate answers when the exact ones are not available. A CQA system also allows conceptual queries where concepts can be used in place of actual values in a query. These capabilities of a CQA system are made possible by using Type Abstraction Hierarchy (TAH)--a knowledge structure that organizes database objects in an IS-A hierarchy where each node represents a concept and approximate answers can be derived by generalization and specialization operations. TAH can be manually generated for small or common sense domains. For large or unfamiliar domains, it is extremely difficult and time consuming to manually generate TAH and, therefore, it is necessary to rely on machine for TAH generation. In this dissertation, we thoroughly investigate the problem of automatic TAH generation. The most important task for automatic TAH generation is to develop a quality measure for TAH, without which it is impossible for a machine to select among numerous possibilities the optimal TAH to generate. Since the major application of TAH is to derive approximate answers, we propose a quality measure for approximate answers called relaxation error. Based on the minimization of relaxation error, efficient algorithms for TAH generation are then developed. Next we generalize the methodology of TAH generation from a single attribute to multiple attributes such that inter-relationship among attributes can be better captured. Finally, we further generalize TAH generation to temporal data. We propose new distance function for time sequences which takes into consideration several important factors that are necessary for deriving good approximate answers but are not accounted for by the existing work. We have implemented the proposed algorithms for TAH generation. Empirical experiments were performed to evaluate the algorithms and the generated TAHs based on a large transportation database. The results are reported in this dissertation which show that the algorithms are very efficient and the generated TAH can be used to derive good approximate answers.