Enriched thesauri and their uses in information retrieval and storage

Michiel Hazewinkel · Centrum Wiskunde & Informatica (CWI), the national research institute for mathematics and computer science in the Netherlands · 1996

The subject of this discussion paper is information storage and, especially, information retrieval from large and very large collections of objects.The focus is on scientific objects such as papers, tables, programs, handbooks, manuals, ....All of these will be referred to here and below as documents.They can be quite heterogeneous in form and content and the storage medium can be varied, though it is assumed that at least adequate metadata (see below for this concept) are, or will be, attached in machine readable form.Quite apart from the sheer size of the problem there are a variety of reasons (for instance linguistic ones having to do with morphological variations, synonyms, and homonyms) that indicate-personally I would put it much stronger-that full text search is not a real alternative.This has always been the major reason to work with a "controlled vocabulary", here interpreted as a standard list of key phrases for a given section of science and/or technology, [2].Another main reason is multilinguality.It is feasible to have essentially the same thesaurus (of key PHRASES (possibly with additional identifiers to deal with ambiguities)) in several languages with good explicit correspondences.It is in any case far simpler to realize such a thing than to do anything like automatic translation. Metadata.The key to dealing with (very) large collections of scientific documents is metadata.This concept includes such things as bibliographic data such as CIP data and Library of Congress cataloguing data; it also includes classifications in terms of one of more (standard) classification schemes and attaching a number of key phrases to a document.The importance of metadata is illustrated, e.g., by the very substantial effort that Elsevier puts continuously into maintaining the thesaurus behind EMBASE, the database of Excerpta Medica, [4], and the even more considerable effort involved in adding adequate metadata in the form of key phrases and words from that thesaurus to each document of EMBASE/Excerpta Medica.

Read the paper · More papers on PaperTik