Language Specific and Topic Focused Web Crawling
Olena Medelyan, Stefan Schulz, Jan Paetzold, Michael Poprat, Kornél Markó · 2006
The Web has been successfully explored as training and test corpus for a variety of NLP tasks ([8], [2], [6]). However, corpora derived from the Web are usually inconsistent and highly heterogeneuos in their nature, which is normally counterbalanced by extending their size to billions of words. We assume that web crawling that takes into account