Delexicalised Multilingual Discourse Segmentation for DISRPT 2021 and Tense, Mood, Voice and Modality Tagging for 11 Languages

Tillmann Dönicke · 2021

This paper describes our participating system for the Shared Task on Discourse Segmentation and Connective Identification across Formalisms and Languages.Key features of the presented approach are the formulation as a clause-level classification task, a languageindependent feature inventory based on Universal Dependencies grammar, and compositeverb-form analysis.The achieved F1 is 92% for German and English and lower for other languages.The paper also presents a clauselevel tagger for grammatical tense, aspect, mood, voice and modality in 11 languages. DatasetSents Conn.Delex.WO deu.rst.pcc2,193 no no OV eng.pdtb.pdtb48,630 yes yes VO eng.rst.gum8,292 no yes -"eng.rst.rstdt8,318 no yes -"eng.sdrt.stac11,087 no no -"eus.rst.ert2,380 no no OV fas.rst.prstc2,179 no no OV fra.sdrt.annodis1,507 no no VO nld.rst.nldt1,651 no no OV por.rst.cstn2,221 no no VO rus.rst.rrt23,044 no no VO spa.rst.rststb2,089 no no VO spa.rst.sctb516 no no -"tur.pdtb.cdtb31,197 yes yes OV zho.pdtb.cdtb2,891 yes yes VO zho.rst.sctb580 no no -"-

Read the paper · More papers on PaperTik