Browsing and searching compressed documents
Raymond Wan · 2003
ii Compression and information retrieval are two areas of document management that exist separately due to the conflicting methods of achieving their goals. This research examines a mechanism which provides lossless compression and phrase-based browsing and searching of large document collections. The framework for the investi-gation is an existing off-line dictionary-based compression algorithm. An analysis of the algorithm, supported by previous work and experiments, high-lights two factors that are important for retrieval: efficient decoding, and a separate dictionary stream. However, three areas of improvement are necessary, prior to the inclusion of the algorithm into a browsing system. First, in order to accommodate retrieval, the algorithm must produce a dictionary built up on words, rather than characters. A pre-processing stage is introduced which separates the message into words and non-words, along with word modifiers. Second, the memory requirements of the algorithm prevent the processing of large