Adaptive Subject Description in Document Retrieval.

Michael D. Gordon · Deep Blue (University of Michigan) · 1984

The central problem in document retrieval is that the subject of a document may be described in many different ways and , similarly, different inquirers may express similar information needs by a variety of different queries. This variance makes it difficult to get the "right" documents into the h and s of the "right" inquirers, for retrieving a document by means of its subject description depends on that subject description adequately matching an inquirer's query. Document descriptions comprise only one part of a retrieval system, and a "good" document description is one that describes the subject of a document in a way that will match the queries of inquirers who will find that document relevant to their information need. In this study, we see that communication (feedback) about the queries of inquirers searching for a given document can be incorporated by a retrieval system in order to redescribe that document so that its description matches better those queries. An adaptive (genetic) algorithm, responsible for such redescription, achieves two aims: first, it increases the probability of a document's subject description matching a query to which the document is relevant (equivalently, it increases the degree of association between a document and a relevant query); second, the algorithm decreases the probability of a document's subject description matching a query to which the document is not relevant (equivalently, it decreases the degree of association between a document and a non-relevant query). Simulation experiments demonstrate the success of adaptive subject redescription in achieving these aims. The simulation technique, itself, is novel: By establishing a set of queries, (to some of which a document is relevant, the rest of which it is not), and measuring the association between the document's description and each of these queries, we obtain estimates of system recall and fallout without building an actual document collection. The method of obtaining such "simulated queries" is described. The simulation technique may help provide a solution to the problem of predicting the performance of a large-scale retrieval system based on its operation in a smaller-scale experimental setting.

Read the paper · More papers on PaperTik