Building a Discourse-annotated Dutch Text Corpus
N.H. van der Vliet, I. Berzlánovich, Gosse Bouma, Markus Egg, Gisela Redeker, S. Dipper, H. Zinsmeister · 2011
We are compiling a corpus of Dutch texts annotated with discourse structure and lexical cohesion, containing initially 80 texts from expository and persuasive genres. We are using this resource for corpus-based studies of discourse relations, discourse markers, cohesion, and genre differences. We are also exploring the possibilities of automatic text segmentation and semi-automatic discourse annotation. This paper discusses our design choices in text selection and segmentation and in the annotation of discourse structure and lexical cohesion. 1