Information retrieval by text skimming
Michael L. Mauldin · Medical Entomology and Zoology · 1989
Most information retrieval systems today use only the presence or absence of keywords to classify and retrieve texts. But simple word searches and frequency distributions do not provide these systems with any understanding of those texts. We call this limit the keyword barrier. To go beyond this barrier, information retrieval systems must at least partially understand the texts they retrieve. Although full natural language parsers are capable today of deep understanding within limited domains, they are still too restrictive and slow for general information retrieval. Text skimming parsers, such as DeJong's F scRUMP, are capable of coarse-level understanding, but they require large amounts of domain-specific knowledge in each application domain. This dissertation describes F scERRET: a full text, conceptual information retrieval system that uses a partial understanding of its texts to provide greater precision and recall performance than keyword search techniques. F scERRET parses its input documents by text skimming and then stores their representations as canonical case frames (called abstracts). User queries are similarly converted to case frames, and are matched to the abstracts using a case frame matcher. F scERRET uses the M sc CF scRUMP parser, a derivative of F scRUMP with two important additions. First, it is able to access an on-line English dictionary (Webster's Seventh) to handle unknown words, using script-based expectations to resolve multiple-meaning ambiguities. Second, a script learning component based on Holland's genetic algorithms updates the script database, augmenting M sc CF scRUMP's episodic world knowledge. Comparison studies of F scERRET's retrieval performance on 1065 astronomy texts show significant improvement in both recall and precision versus the standard boolean keyword search. Precision increased from 35 to 48 percent, and recall more than doubled, from 19 to 52 percent. The script learning component generated new scripts that significantly improved the recall performance of the basic F scERRET system without significant effects on precision. The robust parsing abilities demonstrated, with the depth and flexibility provided by on-line dictionary access and script learning, make text skimming a useful foundation for many applications. The partial understanding provided by a canonical case frame representation is useful for tasks as diverse as information filtering, routing, categorization and summarization.