Knowledge Navigation Librarians in the Word Fray

Karen Hovde · Bulletin of the American Society for Information Science and Technology · 1996

There are 616,500 word-forms in the English language, according to the Oxford English Dictionary, the 20 volumes of which explicate and preserve the richness of our language. Many familiar linguistic arenas—film and television broadcasting, political rhetoric, children's literature—are undergoing a decline in the breadth and complexity of vocabulary usage. In the information business, however, words are thriving; a deluge likened to the oysters in "The Walrus and the Carpenter," "And thick and fast they came at last, and more, and more, and more—" (Lewis Carroll) Most adults take their words for granted, secure in the effectiveness of their vocabularies to sustain and facilitate communication. The new communications technologies are breaching that dike of language complacency. No other single invention since the printing press has had the same intimate and frustrating relationship with words as the computer. Although computers began their existence as computing devices for numerical and mechanical problems, they proved to be remarkably proficient at word manipulation. The recent explosion of computer software and products, such as CD-ROM and online databases, overwhelmingly has to do with words. We don't call it word manipulation; we call it information, the handmaiden of knowledge. But the organization, the communication of, the search for, that knowledge—in fact, the entire interface between the "stuff" we want and the methods for retrieval—is composed of words. Our ability to identify and extract information from a database is directly related to our skill in manipulating those words. Human language acquisition is a natural process under normal conditions. It may be retarded or even, under tragic circumstances, completely inhibited, but the great majority of children acquire a reasonable level of language competence without directed intervention. For example, the passive vocabulary of most native speakers at the time they enter grade school is some 4000 words, which comprise approximately 1000 "lemmas" (the root form and its variants). Tom Schachtman, writing in The Inarticulate Society, says those 1000 lemmas constitute the cluster of most commonly used words in the language. Vocabulary learning, of both words and word types, progresses within the context of formal schooling. By third grade, children usually recognize 9000 words, reading and understanding 3000 of them, and it is at this time that the shift from reading for word recognition to reading for content takes place. Toward the other end of the continuum, an average number of different words known by a college student is 16,700. These figures, however, reflect basic identification of probable meaning. It does not mean that 16,000 different words are actively and/or accurately used, nor does it indicate classification of words by frequency use (i.e., common or rare words). Language research on American English is in fact revealing trends toward an overall loss over the last 50 years in the range of words used. Vocabularies seem to be diminishing. Without straying into the controversy surrounding living language or cultural literacy, there are implications of this vocabulary loss for information literacy. Multiple definitions of information literacy have been generated by librarians and information scientists, most of them reflecting an academic concern for individual fulfillment and the guarantee of the survival of democracy. According to Shirley Behrens, these definitions, translated into more prosaic terms, share an emphasis on the delineation of an information transfer process, which is composed of information resources and the techniques and skills needed for access and evaluation. Information vendors are interested in the resources themselves and in the creation of products for which end-users will pay; educators focus on the potential end-value of the transfer (i.e., additional knowledge). However, it falls to librarians and other information specialists to deal with the transfer process itself. They have willingly but unwittingly embraced a role, in which they have little power but a great deal of responsibility, as navigators. It is at this point that we return to the conjunction of computers and words. It is not the librarian's task to teach vocabulary. It would not, in the 1960s or 1970s, have been the librarian's task to teach computers. Most librarians would agree with Behrens that it is currently one of their tasks to teach the "processes for acquiring information, including systems for information identification and delivery." Evaluation of the information channel and/or the information delivered may or may not be included in this process. The channel or medium for acquiring information which at present dominates information transfer is the computer, and it imposes certain rules concerning word usage which drive the location of information units (records). The concepts of natural language versus controlled vocabularies as they relate to searching are elementary to the library field, which pioneered the classification of written material by controlled subject headings. The advent of computer-driven search technology has created a new imperative to teach not one, but multiple vocabulary approaches. The systems that bestow the welcome advantages of greatly increased speed, size and definition of searches have by the very volume of the material in the databases necessitated increasingly complex methods of organization and retrieval. This proliferation of controlled vocabularies, which after all are simply an extension and elaboration of natural language categories, complicates the librarian's role because it comes at a time when the general population exhibits less language flexibility than in former decades. An effective vocabulary overlap between the search engine (the indexers) and the user cannot be assumed to exist. And where that overlap is in fact lacking, or minimal, it cannot be completely compensated for by a simple instruction to "use the thesaurus," or "look up the subject headings." The daily struggle on the part of librarians to negotiate the word maze standing between the user and the data he or she seeks is made manifest by the significant percentage of pages in the library literature devoted to the trade secrets of searching. In bibliographic instruction, cataloguing and classification, discipline-specific reference, and computer use in libraries, one finds articles on how to identify, explicate, demonstrate and teach idiosyncratic searching on a multiplicity of databases. What we find in effect are a large number of miniature rutters, the sea-charts kept by the first explorers to successfully traverse sectors of unknown seas. Rutters were not maps. They indicated a single tried path, not the only, or even the best of possible directions—a situation analogous to the lists of Internet addresses passed from individual to individual today. The need for guidance is imposed by the parameters of two fundamental search techniques. Free-text searching: Free-text searching, whether on online catalog systems or subject databases, provides a wider range of terms available for searching. The searcher is not constrained by the use of established terms and headings assigned to the topics, terms he or she may not know. But neither are the chosen words under any constraint within the system. The words are not tied to context. They may be so exhaustive (e.g., war, freedom, control) as to be useless, or too precise (lie detector, not polygraph; date, not acquaintance, rape) to retrieve all desired records. In title searches, where one may search for title words rather than exact titles, there are difficulties with the original word choice on the part of the author. Which synonyms are used? Were the main topic terms used at all in the title and/or abstract? Finally, there is the issue of proximity. As Betty Eddison and David Batty have noted, the semantic association of words within a record as recalled by the "Boolean intersection of natural language terms can produce imprecise results, when two or more terms are present in a record, but without the semantic association expected by the user." Controlled vocabulary searching: Controlled vocabularies and controlled index languages can circumvent some of the problems associated with free-text searching. Context may be defined, synonymous relationships established, and hierarchical organization addressed—whether it be through the mechanism of a thesaurus, the Library of Congress Subject Headings or a preset default structure in which "free" search terms are in fact linked with Boolean connectors to built-in subject headings (WilsonDisc products). However, subject heading searching is not easy, a fact to which those numerous articles on specialized searching stand in mute testimony. In online searches, researchers have noted that as many as one-third of the subject queries users enter into online catalogs fail to produce retrievals. On the other hand, searches can be too successful, producing discouragingly large numbers of records, a phenomenon which then necessitates further manipulation of headings or fields to narrow the search. Expert searchers know that successful searches on any system demand an initial analysis of the database, the search and the search goals. With what kind of controlled vocabulary does the database operate, and how exhaustive is it? Medical databases, for example, may employ many times the number of headings of those found in a social science product. Is the database the equivalent of a print index, with many years of subject heading and index history, or is it a more recent product which covers a smaller field than an entire academic discipline? Is the purpose of the search to find a dozen relevant articles and books or to undertake a full-fledged literature review? Experts use a combination of controlled vocabulary and free search techniques. Where controlled vocabulary is not available, or where it does not yield expected results, searchers move to a free search to corral the relevant vocabulary. Even within the confines of a controlled vocabulary search, free searches may yield results closer to user expectation than would a controlled search. It may similarly be necessary to resort to free-text searching when key search terms are too new or incidental to have been included in the controlled vocabulary. Information vendors seek technological answers to improve the fit between natural-language inquiries and search mechanisms based on controlled vocabularies. PsycLIT is now available, for example, on an interface called Knowledge Finder (from Aries Systems Corporation), which adds natural-language querying and relevance ranking to the range of capabilities from which subscribers can choose. This device takes a natural language phrase, drops out stop words and anything unlikely to contain subject relevance, automatically assigns truncation and then uses a process which involves a probabilistic relevance-ranking algorithm which incorporates fuzzy logic to search for the remaining terms as subject headings. While this kind of system may well be able to "translate" natural language phrases into relevant subject headings, it does not solve the problems of synonymy or hierarchical organization. Searchers for whom comprehensive retrieval is important are directed to use the standard Boolean mode and thesaurus features of PsycLIT. Projects aimed at improving subject access to online catalogs have experimented with grouping items into subject clusters, an automatic "see-also" reference process. Natural language terms are linked to the controlled vocabulary of the LCSH. This procedure does improve search effectiveness, but mapping of the entire LCSH would be a Herculean task and one dependent on search engine and expert system technologies not presently available. When Knowledge Finder retrieves records, it performs the information transfer process, but does not address the process itself. A searcher ought to be able to carry away from the interaction, in addition to the printout of citations, some clue to the organizational and retrieval principles of the search process. Since in all likelihood the user's subsequent return to the system will involve a different search, the specifics of search language become less important than general language principles and search structure. Novice users of these systems are not experts, nor in most instances do they wish to become experts. While the rutter provides valuable direction for expert searchers, it is unwieldy as an instructional device for beginners. An approach which utilizes simple examples to illustrate word usage in information organization and retrieval can give students an introduction to the principles of searching without overwhelming them with the myriad of details involved in searching specific databases. Synonymy: The concept of synonymy seems at times to have been left behind with other impedimenta of secondary school. Studies of college students' abilities to make relatedness judgments on the basis of antonyms and synonyms found that students were more comfortable (and faster) with the former than with the latter. Synonyms, however, can be enormously useful in both free-text and controlled vocabulary searches. The use of alternative words mitigates some of the difficulties of exclusivity and inclusion. It is an easy matter to suggest to students that they be aware that a database may be using adolescents for teenagers or females for women. Examples such as sex role = gender, or results = outcomes alert them to the desirability of trying different words, especially if those first chosen fail to retrieve adequate records. Paraphrasing: Paraphrasing is the foundation of a linguistic analysis technique which takes a title or subject phrase, drops out incidental words, evaluates the relative weight of remaining words and assigns them to search term clusters (the Aries Knowledge Finder approach). Paraphrasing is such a commonplace event in library activities that librarians forget that other people may not automatically evaluate their searches in this way. While more formal paraphrasing (of titles, for example) is effectively incorporated into instruction for library research, it is equally useful with regard to subject headings. A search for American Indians may need to include Native Americans and individual tribal names. Drug-use equates with substance abuse, and the individual who searches for the effects of quitting smoking needs to be made aware that the database probably has records filed away under nicotine addiction or tobacco use. Students are not so much unaware of the alternate terms as they are of the need to try the alternatives. Hierarchical subject organization: Hierarchical organization of subjects and subject term clusters is so embedded in information organization that we underrate its value as an instruction device. Both subject and key word, free-text searches are prone to failure when searchers neglect broad, narrow and related terms. Library of Congress Subject Headings are routinely taught for online catalog searches, but the tenets of hierarchical subject organizations hold true for other databases as well. When Jesse James fails to retrieve records, works on outlaws will include individuals. If kimonos or samurai are too narrow, then Japan—costume, or Japan—history should succeed. Where social insects is too general, or retrieves too many records, species names (bees, ants) may prove more manageable. Individuals do not themselves need to know all of the possible term variants and clusters for a particular search. Even a superficial understanding of how the material is organized can be useful in planning its retrieval. Since words confound the information transfer process, it should be to words that we return to effect a solution. We do not need to teach extensive vocabularies in order to demonstrate the effectiveness of controlled vocabulary searches. One well-chosen word can be sufficient. Use a word like stress or evaluation in quick, comparative free-text and controlled vocabulary searches to show how the records retrieved in each context differ. If we were to teach beginners to approach searches with flexibility—both in the matter of word choice and in the ability to try both free-text and subject searching—they would be equipped with a strategy in which they might place some reasonable expectation of success on multiple systems. Computers are the main information channel in contemporary libraries. Until systems research and technology can devise products which provide easy, universal translation of natural language queries, librarians will struggle daily with the task of being Knowledge Finders. Information literacy and search ability might be better served if we retreat from an emphasis on technique and concentrate instead on the navigational principles of language use.

Read the paper · More papers on PaperTik