LEARNING POS TAGGING FROM A TAGGED MACEDONIAN TEXT CORPUS

Viktor Vojnovski, Tomaž Erjavec · 2005

This paper presents several new linguistic resources for the Macedonian language, in particular a language corpus consisting of the digitized and annotated Orwell's “1984 ” in the Macedonian translation. The produced resources (morphosyntactic specification, lexicon, and corpus) are compatible with the multilingual MULTEXT-East data set. The paper presents the digitisation, up-conversion, alignment, and annotation of the corpus, and then discusses an initial experiment in training and evaluating a Part-of-Speech tagger for the Macedonian language on the produced corpus. 1

Read the paper · More papers on PaperTik