Corpus-based metrics for assessing communal common ground - eScholarship
Roman Kutlák, Kees van Deemter, Chris Mellish · Proceedings of the Annual Meeting of the Cognitive Science Society · 2012
Corpus-based metrics for assessing communal common ground Roman Kutlak ([email protected]) Kees van Deemter ([email protected]) Chris Mellish ([email protected]) Computing Science Department, University of Aberdeen Aberdeen AB24 3UE, Scotland, UK Abstract This article presents the first attempt to construct a computa- tional model of common ground. Four corpus-based metrics are presented that estimate what facts are likely to be in com- mon ground. The proposed metrics were evaluated in an ex- periment with human participants, focussing on a domain of famous people. The results are encouraging: two of the pro- posed metrics achieved a large positive correlation between the estimates of how widely known a property of a famous person is and the percentage of participants who knew the correspond- ing property. Keywords: Common Ground; Common Knowledge; Mutual Knowledge; Evaluation with human subjects; Web as corpus To the best of our knowledge, no general computational models exist for assessing what knowledge is likely to be known. In this paper, we examine a corpus-based strategy for building such a computational model. But, before we go into the details of our approach, there are some terminologi- cal and conceptual issues to be clarified. Common and mutual knowledge have been defined in different ways. In this paper, we shall follow the terminology of Vanderschraaf and Sillari, which the authors clarified with the following example. Suppose each student arrives for a class meeting know- ing that the instructor will be late. That the instructor will be late is mutual knowledge, but each student might think only she knows the instructor will be late. How- ever, if one of the students says openly, “Peter told me he will be late again”, then the mutually known fact is now commonly known. Vanderschraaf and Sillari (2009) Introduction Assessing other people’s knowledge is crucial in many sit- uations. Teachers, for example, do well to highlight infor- mation that their pupils do not know. Examples in other ar- eas abound. Suppose, for example, we want to persuade you to reduce your intake of butter. We might do this by telling you ”butter gives you high cholesterol”. This argument only works if you, the hearer, know that cholesterol is bad for you, as is often assumed, for instance because it raises the likeli- hood of heart disease. The (presumed) fact that cholesterol is bad for you happens to be well publicised, and this might be what lies behind our assumption that you know it. Similar examples obtain in advertising, where companies might per- suade you to buy a toothpaste by saying it contains fluoride, because they assume that many viewers know that fluoride is good for your teeth. It is often important to distinguish be- tween knowledge and belief, but we will focus on cases where the distinction is less than crucial. The difference between information assumed to be “given” (i.e., known by the hearer) and “new” (i.e., privileged infor- mation of the speaker) is crucial to philosophers, logicians and linguists (Frege (1892 (1952)); Strawson (1952); Van Ei- jck (1993), to mention but a few) and it is highly relevant to computational linguists working on Natural Language Gener- ation (NLG) programs (Reiter & Dale, 2000), whose output is meant to mimic human language use. A central example is the generation of referring expressions, which has been studied extensively over the last 20 years (Krahmer & van Deemter, 2012). For example, an NLG program that aims to identify a person would do well to express properties that are likely to be known by the reader. For example, the expres- sion “the former member of Led Zeppelin” would not be very informative to a hearer who has never heard of Led Zeppelin. Thus, mutual knowledge is knowledge shared by a group of people. Common knowledge might be informally charac- terised as knowledge that is publicly shared by a group of people. Slightly more precisely, A and B have mutual knowl- edge of p if and only if A knows p and B knows p. They have common knowledge of p if they have mutual knowledge of p, and A knows that B knows p, and B knows that A knows p, and A knows that B knows that A knows p, and so on, ad infinitum (Lewis, 1969). Logicians and game theorists have proposed various precise definitions of common knowledge (including cases with more than two knowers), typically cast in epistemic logic, which formalise the “ad infinitum” (above) in different ways (Vanderschraaf and Sillari (2009)). For rea- sons that will become clear later, we use a third term that is of- ten used in this connection, common ground, in a loose sense, when the distinction between mutual and common knowledge is irrelevant. The psychologists Clark and Marshall observed that, in simple situations, common knowledge is enforced by “triple co-presence”, where the speaker, the hearers and entities are physically present and the speaker believes that the hearers attend to the entities (Clark and Marshall (1981)). They con- trast this simple situation (which they call personal common ground) with communal common ground, which arises not from physical co-presence but from being in a shared com- munity (e.g., people living in Paris). Speakers are frequently able to distinguish between knowledge that is available to