The Semantic Distance Model of Relevance Assessment.
Terrence A. Brooks · 1998
This paper presents the Semantic Distance Model of Relevance Assessment. It is a cognitive model of the relationship between semantic distance and relevance assessment. Premises of the model such as the subjective nature of information and the metaphor of semantic distance are discussed. Empirical results illustrate the effects of semantic distance and semantic direction. Relevance auras, a combination of vertical and horizontal relevance assessments are also presented. An ongoing series of experiments (Brooks, 1995a, 1995b, 1997) has investigated the relationship between semantic distance and relevance assessments. This essay cumulates these results and adds some new experimental results in presenting the Semantic Distance Model (SDM) of Relevance Assessment. The SDM is a cognitive model describing the relevance relationships between bibliographic database records and topical subject descriptors. The SDM proposes that relevance assessment systematically declines with semantic distance. Such a decline of stimuli effect as a factor of distance in psychological space is typical of many cognitive models (Shepard, 1987). The SDM provides empirical support for the general indexing/cataloging notion that subject-term hierarchies express a conceptual arrangement from broad to narrow. The SDM also provides empirical support for the notion that the “appropriate” subject descriptor will be perceived as more relevant than those at greater semantic distance. A semantically distant descriptor might be one that resides at either a broader or narrower position in the same subject-term hierarchy. A number of unanticipated findings have also been observed: (1) It appears that the SDM can be switched off and on depending on the context of relevance assessment, (2) It appears that topical subject expertise steepens the rate of decline of relevance assessment, and (3) It appears that relevance assessment is contingent on semantic direction; that is, the rate of decline differs if one is assessing a series of increasingly narrow descriptors, or increasingly broad descriptors. The archetype trial of the experiments supporting the SDM has been the comparison of two texts. One text has been bibliographic records trimmed to title and abstract, while the other text has been topical subject descriptors. Relevant assessments were harvested from library science, economics, engineering and mathematics graduate students as they assessed the relevance relationship between subject descriptors and trimmed bibliographic records from the literatures of education, economics, engineering and mathematics. INFORMATION AS A SUBJECTIVE PHENOMENON A fundamental assumption of the SDM is that information is a subjective phenomenon, a product of the interpretation of texts (Cole, 1994; Dervin & Nilan, 1986). This assumption is widely supported. The transactional model of reading asserts that the active negotiation between readers and texts produces meaning (Straw, 1990). Jacob and Albrechtsen (1997, p. 42) describe reading “not a passive reception of meaning unpacked from a neutral and fixed, or stable, reference but an active process of contextualizing the word within the particular set of relativized and competing definitions that inform the moment of expression”. Bertrand-Gastaldy, et al. (1995, p. 56) consider text as an interweave of multiple semiotic systems: “It is not the character strings that are significant; rather, it is their properties specific to each of these systems and interpreted by a cognitive agent.” Since information is a subjective phenomenon, it is hardly surprising that readers will disagree over the meanings of text. The disagreement exhibited by indexers and catalogers has already been widely documented (Broadbent & Broadbent, 1978; Thorngate & Hotta, 1990). There appear to be a number of cognitive factors at work. The frequency distribution of names for objects prompted Furnas, et al. to report that “no single name, no matter how well chosen, can do very well” (1983, p. 1,796). Some of the variability of naming is due to fuzzy natural categories that permit verbal concepts to be modified to suit changing situations (McCloskey & Clucksberg, 1978); other variability is due to the graded structure of categories that permits category members to be ranked according to the category ideal (Rosch, 1975; Shoben, 1976). Recognizing the subjective nature of information does not deny the possibility of agreement about the meaning(s) of a text. If one were to control factors like vocabulary, culture, subject expertise, and so on, a group of readers may agree about the meaning of a text. The work of indexers/catalogers is premised on this potential agreement. It also gives force to the concept of topical relevance. Analogous agreement underlies the basic category of abstraction of ordinary objects: We will argue that categories within taxonomies of concrete objects are structured such that there is generally one level of abstraction at which the most basic category cuts can be made. In general, the basic level of abstraction in a taxonomy is the level at which categories carry the most information, possess the highest cue validity, and are, thus, the most differentiated from one another. (Rosch, et al., 1976, p.382) Online searching is one battleground where the subjective nature of information and textual core meanings clash. The archetypal problem of online searching is disagreeing with the indexer’s labeling of a record; or put another way, being surprised by the records fished out of a database by a given descriptor. Relevance assessment research evolves by investigating textual meanings under controlled conditions. The SDM can be viewed as a device for calibrating agreement about the relationship of bibliographic records and subject descriptors. It describes how the average reader will assess relevance relationships. It does not predict how any particular reader will assess the relationship between a specific subject descriptor and bibliographic record. THE METAPHOR OF SEMANTIC DISTANCE Semantic distance is a psychological construct that has been used to locate concepts along various dimensions of meaning (Schvaneveldt, Durso & Mukherji, 1982). Placing database records in ndimensional space, or docuverse (Das Neves, 1997, Wallmannsberger, 1991) has been an ancient and seductive information conceit. As early as 1973, Charles Bachman urged database programmers to forsake their linear record model in favor of n-dimensional data spaces. Twenty five years later, the spatial metaphor is still a mainstay of theorizing. Chen et al, (1998), for example, proposed a concept space approach to internet searching. The SDM used a variety well-known trees of descriptors and established hierarchies of subject terms, as well as non-library hierarchies, to provide semantic distance. The witness of the SDM with a variety of hierarchies, including non-library ones, suggests its robustness. So far, an SDM effect has been produced by the following subject hierarchies and descriptor trees: * The Thesaurus of ERIC descriptors in Brooks (1995a, 1995b, 1997). * The Index of Economic Articles in Journals and Collective Volumes in Brooks (1995b). * Mathematical Reviews in Brooks (1995b). * INSPEC Thesaurus, 1993. The Institution of Electrical Engineers. (As yet unpublished, but a portion of the results appears below.) * Thesaurus of Psychological