Indexing and automatic significance analysis

Ivo Steinacker · Journal of the American Society for Information Science · 1974

Abstract Intellectual indexing proceeds on three levels: The selection of phrases occurring in the document text (sequential indexing), the posting of specific phrases from the text to generic descriptors (generic indexing), and the choice of descriptors which are implicit to the document text (symbolic indexing). Automation has been attempted on all three levels: by concordance and autoposting. Here an algorithm is proposed to solve the problem of sequential indexing which does not use any grammatical or semantic analysis, but follows the principle of emulating human judgement by evaluation of machine‐recognizable attributes of structured word assemblies (text). The algorithm is based on producing “text cuts” of a few words in length and ordering them alphabetically. Afterwards, every “text cut” which appears with a certain limit frequency or above is considered significant (by human standards). The algorithm has been applied to a text body of about 220,000 words from the NASA bibliographic file and an “established” dictionary of significant terms has been created by this algorithm. As any phrase not occurring in the established dictionary is not suppressed, but posted to a floating dictionary, from which it may, if usage increases above the limit frequency, be transferred to the established dictionary, the algorithm presents a tool for the creation and maintenance of a “self‐adaptive” data base of text information.

Read the paper · More papers on PaperTik