On Bias-free Crawling and Representative Web Corpora
Roland Schäfer · 2016
In this paper, I present a specialized opensource crawler that can be used to obtain bias-reduced samples from the web.First, I briefly discuss the relevance of bias-reduced web corpus sampling for corpus linguistics.Then, I summarize theoretical results that show how commonly used crawling methods obtain highly biased samples from the web.The theoretical part of the paper is followed by a description my feature-complete and stable ClaraX crawler which performs so-called Random Walks, a form of crawling that allows for bias-reduced sampling if combined with methods of post-crawl rejection sampling.Finally, results from two large crawling experiments in the German web are reported.I show that bias reduction is feasible if certain technical and practical hurdles are overcome.