Extracting and Annotating Wikipedia Sub-Domains — Towards a New eScience Community Resource

Gisle Ytrestøl, Dan Flickinger, Stephan Oepen · 2008

We suggest a simple procedure for the extraction of Wikipedia sub-domains, propose a plain-text (human and machine readable) corpus exchange format, reflect on the interactions of Wikipedia markup and linguistic analysis, and report initial experimental results in parsing and treebanking a domainspecific sub-set of Wikipedia content.

Read the paper · More papers on PaperTik