Extending relational database management systems for information retrieval applications

Clifford A. Lynch · University of California, Berkeley eBooks · 1987

This thesis studies the use of relational database systems to construct large, high performance information retrieval systems such as online library catalogs or citation retrieval applications. The major problem areas in relational implementations are query execution costs, poor space utilization, and functionality deficiencies both in query processing and in query languages such as SQL. Analytic and simulation methods are applied to quantify these problems. Proposals extending earlier work on user-defined operators for relational query languages and accompanying secondary index support allow both efficient query formulation and the definition of space-efficient relational bibliographic databases. When column values follow distributions typical of bibliographic databases (Zipf distributions), a key performance problem is inaccurate selectivity estimation. A framework for incorporating user-defined selectivity estimators into a relational query optimizer is established, and methods are given to construct highly accurate selectivity estimators for bibliographic databases. Relational query optimizer extensions are specified which incorporate query execution plans that use TID list manipulation algorithms for evaluating single-relation queries into the optimizer's vocabulary. With these extensions a relational system can outperform an inverted file retrieval system on bibliographic databases. Also explored are query planner extensions to implement nonmaterialized relations (allowing both partially deferred evaluation of queries and inexpensive iterative query construction) and preexecution identification of queries that will be costly to evaluate or will produce very large results. Both of these features are important for public access information retrieval applications. Finally, the thesis examines difficulties that arise in using a relational query language to support advanced information retrieval techniques such as ranking and weighted retrieval, and develops query language extensions that would significantly improve the performance of such searching techniques in a relational setting.

Read the paper · More papers on PaperTik