Efficient Transaction Management and Query Processing in Massive Digital Databases
Mohan Kamath, Krithivasan Ramamritham · 1995
We address several important issues that arise in the development of Massive Digital Database Systems (MDDSs) in which data is being added continuously and on which users pose queries on the fly. {\em News-on-demand} and document retrieval systems are examples of systems that have these characteristics. Given the size of data, metadata such as index structures become even more important in these systems --- data is accessed only after processing the metadata, both of which will reside on tertiary storage. The focus of this paper is on {\em query and transaction processing} in such systems, with emphasis on {\em metadata management}. The performance in these systems can be measured in terms of the {\em response time} for the queries and the {\em recency} or age of the items retrieved. Both need to be minimized. The key to satisfying the performance requirements is to exploit the characteristics of the metadata as well as of the queries and updates that access the metadata. After analyzing the functionality and correctness properties of updates, we develop an efficient scheme for executing queries concurrently with updates such that the queries have short response times and are guaranteed to return the most recent articles. Secondly, we address logging and recovery issues and propose techniques for efficiently migrating metadata updates from disk to tape. Thirdly, considering the tape access needs of queries, we develop new tape scheduling techniques for multiple queries such that the response time of queries is reduced. Results of the performance tests on a prototype system show the superior performance of the developed algorithms and reveal that to build high performance MDDSs it is imperative that we adopt approaches that exploit the data and transaction characteristics.