A comparable Wikipedia corpus: from wiki syntax to POS tagged XML

Noah Bubenhofer, Stefanie Haupt, Horst Schwinn · Publication Server of the Institute for German Language (Institute for German Language) · 2016

To build a comparable Wikipedia corpus of German, French, Italian, Norwegian, Polish and Hungarian for contrastive grammar research, we used a set of XSLT stylesheets to transform the mediawiki anntations to XML. Furthermore, the data has been amnntated with word class information using different taggers. The outcome is a corpus with rich meta data and linguistic annotation that can be used for multilingual research in various linguistic topics.

Read the paper · More papers on PaperTik