CommonCOW: Massively Huge Web Corpora from CommonCrawl Data and a Method to Distribute them Freely under Restrictive EU Copyright Laws
Roland Schäfer · 2016
In this paper, I describe a method of creating massively huge web corpora from the CommonCrawl data sets and redistributing the resulting annotations in a stand-off format.Current EU (and especially German) copyright legislation categorically forbids the redistribution of downloaded material without express prior permission by the authors.Therefore, stand-off annotations or other derivates are the only format in which European researchers (like myself) are allowed to re-distribute the respective corpora.In order to make the full corpora available to the public despite such restrictions, the stand-off format presented here allows anybody to locally reconstruct the full corpora with the least possible computational effort.In Section 1., I briefly introduce the technology behind the COW project (Corpora from the Web), which is used to create the CommonCrawl-derived corpora.In Section 2., I provide some details about the resulting CommonCOW (COCO) web corpora.Finally, in Section 3., I introduce a method to circumvent restrictive EU copyright laws by distributing only the corpus annotations (under a CC-BY license) together with a tool that allows users to locally reconstruct the corpus from the annotations and the original CommonCrawl files.