A Natural Language Information Retrieval System
H.M.G.M. Jacobs · Methods of Information in Medicine · 1968
This paper describes a system for dealing with a certain kind of textual information. The system has been in operation for about one and a half years. It also indicates the nature of a new and greatly expanded system presently under development. The first system consists essentially of three parts, of which the most important is a thesaurus processor. The assumption is made that document content depends only on word content and that word relationships are defined by an hierarchical structure. The function of the thesaurus processor is to provide a simple language for developing and changing the thesaurus, whenever change is necessary. The remaining parts of the first system are a document processor for updating document files, and a search processor for batch requests which must scan the entire document file. Some statistics are given on performance of the system. The second system includes a thesaurus processor of expanded capability. It also includes a newly developed search language, which can be used to scan records for complex patterns of events. Input-output processing of records and of search results is left to the user of the system.