On Bias-free Crawling and Representative Web Corpora

Roland Schäfer · 2016

In this paper, I present a specialized opensource crawler that can be used to obtain bias-reduced samples from the web.First, I briefly discuss the relevance of bias-reduced web corpus sampling for corpus linguistics.Then, I summarize theoretical results that show how commonly used crawling methods obtain highly biased samples from the web.The theoretical part of the paper is followed by a description my feature-complete and stable ClaraX crawler which performs so-called Random Walks, a form of crawling that allows for bias-reduced sampling if combined with methods of post-crawl rejection sampling.Finally, results from two large crawling experiments in the German web are reported.I show that bias reduction is feasible if certain technical and practical hurdles are overcome.

Read the paper · More papers on PaperTik