Annotating Information Structure: The Case of Topic
Philippa Cook, Felix Bildhauer · 2011
paper deals with the annotation of Sentence Topics/Aboutness Topics in naturally occurring data. We report on a corpus study in which relatively poor inter-rater agreement was attained for the annotation of topics, although both coders were adhering to the same annotation instructions. Tokens that were particularly difficult to assess are identified, systematized, and discussed in some detail. In sum, the cases that are most likely to lead to non-matching annotations are those that either require a decision between thetic or topic-comment, or involve a no verlap between Focus and Topic. The findings raise a number of issues that may contribute to the discussion in theoretical linguistics, and they alsomay alert other researchers planninga similarenterprise to some pitfalls they may encounter. 1I ntroduction Research on information structure may serve a twofold purpose: first, information structure constitutes an intriguing area of investigation in its own right, where numer- ous concepts and their interrelations are still in need of further refinement. Second, insights in this field may lead to promising (re-)analyses of linguistic phenomena on the basis of information structure, that is, using information structural constraints in describing phenomena previously accounted for in terms of syntax (e.g. De Kuthy, 2002; Cook and Payne, 2006; Ambridge and Goldberg, 2008; Coo ka nd Orsnes, 2010). In this respect, corpora annotated for information structure are particularly valuable, as they put one in a position to test linguistic analyses that ar eb ased on notions such as topic, focus and givenness. However, not only are these notions used in different ways across different currents of research, but they also cause considerable problems when applied to naturally occurring data by researchers who otherwise agree largely on the definitions of these concepts and who even adhere to the same set of annotation guidelines. In the present paper, we will take a closer look at the annotation of Sentence Top- ics/Aboutness Topics in naturally occurring data. The data we will discuss were ex- tracted from the DeReKo 1 corpus and coded for information structure as part of a study on preferential topic realization in German newspaper texts, that is, a corpus study not initially related to the present work. Section 2 outlines th ec riteria used by the an- notators for identifying Aboutness Topics and relates them to alternative approaches to topic-hood. Section 3 reports the relevant details of the corpus study, including measurements of inter-rater agreement for the annotation o fA boutness Topics. In Sec- tion 4, we identify the type of data that turns out to be particularly difficult to assess 1 http://www.ids-mannheim.de/kl/projekte/korpora/