Part of Speech Annotation of a Turkish-German Code-Switching Corpus

Özlem Çetinoğlu, Çağrı Çöltekin · 2016

In this paper we describe our efforts on POS annotation of a code-switching corpus created from Turkish-German tweets.We use Universal Dependencies (UD) POS tags as our tag set.While the German parts of the corpus employ UD specifications, for the Turkish parts we propose annotation guidelines that adopt UD's language-general rules when it is applicable and adapt its principles to Turkishspecific phenomena when it is not.The resulting corpus has POS annotation of 1029 tweets, which is aligned with existing language identification annotation.

Read the paper · More papers on PaperTik