TOPOLOGICAL STRUCTURES OF INFORMATION RETRIEVAL SYSTEMS

Robert Tien-Wen Chien, F. P. Preparata · 1966

Abs tractThis paper considers the problem of information retrieval from the point of view of graph theory.In this formulation documents are represented as nodes and relationships among the documents are represented by edges.Two types of graphs are introduced, namely the similarity graph which is based on subject-content correlation and the citation graph, which is derived from direct citation linkages among documents.Several distance measures are considered and evaluated with regard to retrieval operations.In retrieval operation a query is presented to the system which describes a profile of the type of documents to be retrieved from the collection.In most systems employing coordinate indexing today the query is given in terms of a set of descriptors or some logical function thereof.For instance, we may ask for all documents that deal with the "decoding" of "Bose-Chandhuri-Hocquenghem Codes" that are published in the "Transactions of IEEE on Information Theory" since "1964," where those terms under quotation signs are descriptors.Another type of retrieval systems are based on citation indexing.In this type of systems citation information among documents is stored in the system.The query is given in terms of specifying accession documents in the network.For instance, one might wish to retrieve all documents citing a document d or one might wish to retrieve all documents that are cited by document d.Retrieval operations based on multi-generation citations are theoretically feasible but so far have not received much attention.In comparing the two popular schemes, citation indexing is easy to instrument but is limited in scope in that it derives information only from existing direct linkages in the document collection.This restriction is reflected in the usual incompleteness of retrieval results when one is interested in searches based on subject content.On the other hand, coordinate indexing works well only if the indexed document collection is relatively homogeneous and the query well-defined.For requests from research scientists the query is always aimed at the 0^section or the union of several narrow and ill-defined disciplines.As a result, the outcome is usually contaminated with large amounts of irrelevant material.

Read the paper · More papers on PaperTik