Automatic indexing and retrieval as a tool to improve information and technology transfer

Harald H. Zimmermann · Publications of the UdS (Saarland University) · 1982

During the last 20 years, linguistic data processing mainly has been seen as a tool to develop linguistic regularities (or detect irregularities) of a given natural language, especially to handle large textual databases (Corpora). A second motivation to use a computer was to test some theories or models of a language system (or a of it) using a simulation program. As a result of both strategies, the Text Analysis has been implemented. At present, a very large lexical database is available to analyse written German morphologically and syntactically. The syntactic parser is able to handle every German sentence with more than 90% correct results. On the other hand, the development of large (textual) databases within different fields (e.g. law, patent specifications, medicine) is increasing rapidly. Therefore, a computer aided indexing system (Computergestutzte Texterschliesung: CTX) has been developed at Regensburg and Saarbrucken University to improve the (even natural language oriented) access to textual data (free text) applying linguistic strategies to information retrieval processes. Main results of feasibility studies, especially in the field of German Patent Documentation, are presented. Introduction The central problem in this lecture is, first of all, neither a linguistic nor a computer-technical one. The starting point is rather the problem of developing and processing knowledge coded in natural language, in a way which makes it possible for persons and social groups to identify it safely, and to use it for solving their own problems, or for making decisions. Knowledge coded in natural language is embodied in books, newspaper articles, reviews, legal texts, judgements, reports, patent specifications, journal articles, notices, and so on. In the following, these texts will be called or units. By now, about 100.000 periodicals are published all over the world. It is estimated that there are more than 5 million specialized publications (essays, books) in a year /1/. The data base Chemical Abstracts covered about 5.5 million document units in 1981, and every year about 500.000 document units are added /2/. Traditionally, the access to this knowledge is structured by bibliographies, review journals, and collections of file cards. In the meantime, (mostly specialized) classification systems and subject catalogues have been developed. By the aid of such systems, the contents of the documents have been made accessible. By classifying them according to these criteria and rather formal features (for instance author, year of publication, place of publication), a sure retrieval becomes possible. These works of analyzing intellectually involve high staff, they need a highly qualified collaborator; but in the end they are, in many respects, unsatisfactory, because the strong reduction (normally to few key words) causes a loss of information which hinders the retrieval of relevant documents. Moreover, this procedure requires a specialist of the system with the function of a go-between when making a This causes, once again, high cost and complicates the access (time delay). The information technology presents by now considered only on the technical side tools which would allow to connect any interested person by telephone line or by special networks with an information system where the facts are stored electronically (on-line access to data bases). Usually, the information coded in natural language is, still today, found in printed form which is not directly accessible to computers. Moreover, the possibilities of storage are in spite of considerable improvements of storage capacities still quite restricted. This is why usually only titles or abstracts of essays or books are stored in these data bases (on the intellectual way, so for instance written by the author himself). Other kinds of (patent specifications, judgements) are usually represented by the relevant parts of the text. Nevertheless, these quantities are already so huge (the Legal Information System JURIS of the Federal Republic of Germany has already stored more than 200.000 judgements and bibliographical abstracts) that a linguist will feel giddy in the view of the quantity of these corpora. Very often the text part of such a document in the data base is only an informational instrument not understood by the system -, that is to say: One can only find it and use it in a retrieval process, when the document has already been identified and called up via other items (author, key word ...). Nevertheless, some information systems provide already a so-called free-text retrieval. One should suppose that, on this way, linguistic findings are used. Anyway, this is (almost) never the case in the present situation. On the contrary, the method is usually as follows (for instance in the case of the systems DIALOG, STAIRS, and DIRS/GRIPS): (a) When processing the document, the textual is diminished by the so-called stop which are stored in a special word list; the remaining words are classified and stored as word forms or chain of characters. A document identification connected with the word form makes sure that a document will be found by a word form of the (free) text. (b) In the retrieval, that is during the search for information, the word form can be cut off on the right side (truncation, e.g. by using the §-sign), so that the beginning of a word is sufficient to identify all documents containing words with this beginning.

Read the paper · More papers on PaperTik