An Empirically Grounded Approach to Extend the Linguistic Coverage and Lexical Diversity of Verbal Probabilities - eScholarship
Chrsitine Engelmann, Udo Hahn · Proceedings of the Annual Meeting of the Cognitive Science Society · 2014
An Empirically Grounded Approach to Extend the Linguistic Coverage and Lexical Diversity of Verbal Probabilities Christine Engelmann Udo Hahn Jena University Language & Information Engineering (J ULIE ) Lab Friedrich-Schiller-Universit¨at Jena Jena, Germany [email protected] [email protected] Abstract Our study introduces a rigorous empirical, corpus-based criterion for the selection of relevant items, thus balancing the variety of word classes and, as a consequence, enlarging the lexical diversity dealt with in this area of research. We also collect preliminary evidence for the impact discourse context has on the properly adjusting verbal probabilities. Linguistic expressions indicating uncertainty of states of knowledge or beliefs, such as “possible” or “might suggest”, are usually dealt with in the psycholinguistic community un- der the heading of ‘verbal probabilities’. Despite a remark- able level of quantitative and experimental rigor, studies deal- ing with this phenomenon suffer from several methodological shortcomings: The selection of items under scrutiny usually lacks empirical justification besides subjective preferences, the items are often investigated in isolation, i.e. without sufficient linguistic context and focus is typically on only few word classes, usually adjectives and adverbs. Our study introduces a rigorous empirical, corpus-based criterion for the selection of relevant items, thus balancing the variety of word classes and, as a consequence, enlarging the lexical diversity dealt with in this area of research. We also collect preliminary evidence for the impact discourse context has on the properly adjusting ver- bal probabilities. Related Work There is a long tradition and a vast amount of literature con- cerned with the translation of verbal into numerical probabil- ities (an extensive discussion is provided by Clark (1990)). The approaches are diverse, including e.g. the assignment of numbers to expressions (Lichtenstein & Newman, 1967; Reagan, Mosteller, & Youtz, 1989; Clarke, Ruffin, Hill, & Beamen, 1992), the assignment of expressions to numbers (Reagan et al., 1989), pair comparison (Budescu & Wall- sten, 1985; Wallsten, Budescu, Rapport, Zwick, & Forsyth, 1986) and rank-ordering (Budescu & Wallsten, 1985). Usu- ally, scales ranging from ‘0’ to ‘1’ or from 0% to 100% prob- ability are employed. Despite the differences in methods re- sults are relatively comparable (Reagan et al., 1989; Clarke et al., 1992). Teigen and Brun (2003) summarize the main findings by stipulating two claims—a high degree of similar- ity in the mean estimates between study groups, on the one hand, and a high degree of inter-individual variability within groups, on the other hand. These observations have led researchers to focus on the inherent vagueness of probability expressions. The core of such investigations is the modeling of verbal probabilities as fuzzy concepts and their subsequent characterization as membership functions over the probability scale (Wallsten & Budescu, 1995). In this respect, probabilities can be as- signed values ranging from ‘0’, if they are not included in the concept, to ‘1’, if they are perfect exemplars of the con- cept. The vagueness of a specific expression is then repre- sented by location, range and shape of the membership func- tion. Recently, this approach has been adapted in a study by Bocklisch, Bocklisch, Baumann, Scholz, and Krems (2010). The authors describe a two-step procedure which includes di- rect estimations from participants of minimal, maximal and best corresponding probability values, as well as data anal- ysis in terms of membership function construction. Further- more, considerable work has been carried out on factors that might influence the interpretation and choice of verbal prob- abilities, e.g. extra-linguistic context (Brun & Teigen, 1988), Keywords: verbal probabilities, epistemic modality, empir- ical semantics, uncertainty in language comprehension Introduction Our daily communication is full of linguistic signals to indi- cate lack of certainty or different degrees of belief in what we are saying. Choices of modal verbs (“may”), adjectives (“possible”), adverbs (“probably’) or lexical verbs (“sug- gest”), etc. are adequate means to calibrate the likeliness we attribute to a proposition we utter. The relevance of this phe- nomenon, commonly called verbal probabilities in the psy- chological community and epistemic modality in the linguis- tic community, has early been recognized by cognitive scien- tists who focus on the study of language comprehension (cf., e.g. Lichtenstein and Newman (1967)). Still, the way these investigations have been carried out up until now suffers from several methodological shortcomings. First, the specific lexical items are collected with a consider- able subjective bias mostly based on individual preferences. To the best of our knowledge, there is no study which justi- fies the selection of items under scrutiny by empirical criteria (e.g. distribution frequencies in a corpus). Given the long his- tory of lexical association tasks in cognitive science, there is also no wonder that verbal probabilities are primarily stud- ied without linguistic context. So many studies focus on the probability of “possible” or “likely” in complete isolation. Finally, the focus of previous work has predominantly been on few selected word classes, such as adjectives and adverbs, without paying equal attention to lexical verbs or nouns as carriers of probability information.