A Parallel Architecture for Query Processing Over A Terabyte of Text

Peter Bailey, Peter Bailey, David Hawking, David Hawking · ANU Open Research (Australian National University) · 1996

The Parallel Document Retrieval Engine (PADRE) has previously demonstrated that full text scanning methods supported by parallel hardware permit powerful query constructors and rapid response to changing document collections. Extensions to PADRE have been designed and implemented which make use of parallel secondary storage to allow each procesing node to handle data up to 32 times the size of its primary memory. Using the largest purchasable machine on which PADRE currently runs, these increase the maximum possible collection size to one terabyte. This paper addresses the practicality of achieving this limit and the extent to which the performance, responsiveness, functionality and scalability of the full text scanning PADRE are preserved in the extended version. KEYWORDS Text retrieval, indexes, dictionaries, parallel computing. 1 Introduction PADRE [8, 5] is a free text system designed to retrieve documents relevant to a specified research topic from among a large collection. Ther...

Read the paper · More papers on PaperTik