Building Linguistic Corpora from Wikipedia Articles and Discussions
Eliza Margaretha, Harald Lüngen · LDV-Forum/Journal for language technology and computational linguistics · 2014
Wikipedia is a valuable resource, useful as a lingustic corpus or a dataset for many kinds of research.We built corpora from Wikipedia articles and talk pages in the I5 format, a TEI customisation used in the German Reference Corpus (Deutsches Referenzkorpus -DeReKo).Our approach is a two-stage conversion combining parsing using the Sweble parser, and transformation using XSLT stylesheets.The conversion approach is able to successfully generate rich and valid corpora regardless of languages.We also introduce a method to segment user contributions in talk pages into postings.