Compression and fast indexing for multi-gigabyte text databases
Alistair Moffat, Justin Zobel · 1994
In the last two years we have developed improved techniques for indexing and retrieval of text data, including algorithms for inversion, for compression of the data and index, and for economical ranking. These techniques were, however, tested on relatively small databases. In this paper we describe our experiences in scaling these techniques up to a large (2 Gb) heterogeneous text database. Our experiments show that compression performance does not degrade with the increase in size, and that response times remain small, confirming that our techniques are suitable for large volumes of data. 1 Introduction Large document collections present many problems for practical information retrieval. To avoid unnecessary accesses to the text of the collection during query evaluation, comprehensive indexes are required. It must also be possible to create and access these indexes in a reasonable amount of time, and to store them, and the data itself, in a reasonable amount of space. In the last two...